Abstract
Background
This study aimed to review level I and II therapeutic studies on boxer’s fractures to measure variation in quality among the highest level study designs.
Methods
We used quantitative measures of study quality to evaluate prospective randomized controlled trials (RCTs) of treatments of boxer’s fractures. A search of PubMed, using terms “boxer’s fracture” and “fifth metacarpal neck fracture” identified 164 articles from 1961 to 2019. From this list, we identified 6 RCTs. Two observers classified each trial according to 3 systems: the Oxford Levels of Evidence, the modified Coleman Methodology Score, and the revised Consolidated Standards of Reporting Trials (CONSORT) score.
Results
The 2 reviewers were consistent in their use of the Oxford Levels of Evidence (100% agreement). The differences between the average modified Coleman Methodology scores and the average CONSORT scores assigned by the 2 observers were not significant (46.2 vs 45.3 points, κ = 0) and (13.7 vs 14.3 points, κ = 0.33), respectively. Both observers rated all the studies as level I and as unsatisfactory according to the Coleman Methodology Score (100% and 100%), and less than half as unsatisfactory according to the CONSORT score (50% and 17%). Areas of deficiency included randomization, blinding, group comparability, clinical effect measurements, and allocation into treatment arms.
Conclusion
Classifying orthopedic scientific reports according to the levels of evidence implies a degree of respect for level I and II studies that may not always be merited. Our data suggest that the quality of higher level studies, namely those involving boxer’s fractures, varies and may often be unsatisfactory when critically evaluated.
Keywords
Introduction
The Oxford Levels of Evidence are commonly assigned to medical publications and tend to carry a large degree of weight when accessing the validity of study conclusions and study design.1-4 However, these conclusions may be problematic as high levels of evidence do not always correlate to a high-quality scientific study. 3 Conversely, studies with lower levels of evidence are often undervalued even when they may be higher quality scientific studies in terms of study design or even clinical impact. 5
Fractures of the fifth metacarpal neck, or boxer’s fractures, are common injuries, especially in the young male population.6-8 Traditionally, boxer’s fractures have been treated nonoperatively with closed reduction and splint immobilization.9,10 However, a recent systematic review of several prospective randomized controlled trials demonstrated that there is no clear benefit to reduction and cast immobilization when compared with soft wrap techniques. 11
The purpose of this study was to review the quality of high-level scientific studies (level I and II prospective randomized controlled trials) in the treatment of boxer’s fractures. We hypothesized that the methods of these randomized controlled trials would be sufficient to support the conclusion that boxer’s fractures do not need reduction and can simply be treated with soft wrap techniques.
Materials and Methods
A search of PubMed, 12 with the use of combinations of the terms “boxer’s fracture” and “fifth metacarpal neck fracture,” identified 164 articles published from 1961 to 2019. From this list, we identified 6 prospective randomized controlled trials that specifically addressed the nonoperative management of boxer’s fractures.13-18 Two observers (both board-certified hand and upper extremity surgeons with MD degrees) classified each trial according to 3 systems: the Oxford Levels of Evidence, 1 the modified Coleman Methodology Score (0-100 point scale),3,19-21 and the revised Consolidated Standards of Reporting Trials (CONSORT) score. 22
Each observer first assigned an Oxford Level of Evidence for each paper. 1 Using this system, level I studies were defined as high-quality randomized controlled trials with greater than 80% follow-up. Level II studies were defined as low-quality studies, such as randomized controlled trials with less than 80% follow-up, or prospective/retrospective cohort/case-control studies without randomization, among other considerations.
Each observer then assigned a modified Coleman Methodology Score to each study.3,19-21 Initially, the Coleman Methodology Score was developed to grade the methodology of clinical studies on patellar and Achilles tendinopathy.19-21 This system was subsequently modified in a quality analysis study of lateral epicondylitis clinical trials to best access the quality of prospective randomized controlled trials. 3 Using this modified Coleman Methodology Score, 15 individual criteria were used to create a solitary score for each study. The total score had a range of 0 to 100 points. The criteria and scoring system are illustrated in Table 1, based on the work performed by Cowan et al. 3 We used the categorical rating system, which defines each study to be “excellent” if the score is 85 to 100 points, “good” for 70 to 84 points, “fair” for 55 to 69 points, and “poor” for less than or equal to 54 points. A rating of “fair” or “poor” was deemed unsatisfactory (Table 1).
Modified Coleman Methodology Score. 3
Finally, each observer then assigned a revised CONSORT score to each paper. The revised CONSORT score is a 22-item checklist that assesses various aspects of prospective randomized trials. 22 The purpose of this checklist is to provide a standardized approach to compare the conduct of trials and the validity of their results. 22 Within the 22-item checklist, the trial is given either 1 point for each criterion met or 0 points if the criteria were not met, with a maximum score of 22 points. The categorical rating system assigns each study to be “excellent” if the score is 18 to 22 points, “good” for 13 to 17 points, “fair” for 8 to 12 points, and “poor” for less than or equal 7 points. A rating of “fair” or “poor” was deemed unsatisfactory (Table 2).
CONSORT Checklist of Items to Include When Reporting a Randomized Trial. 3
Note. CONSORT = Consolidated Standards of Reporting Trials.
For continuous variables, the Student t-test was used to assess the interobserver reliability of each scoring system. For categorical variables, such as those seen in the Modified Coleman and CONSORT scores, the kappa statistic was used to evaluate the interobserver reliability of each scoring system. Kappa scores range from −1 to 1, with 1 indicating perfect agreement and a kappa of 0 indicating agreement equivalent to chance. 23 A categorical rating of the reliability, first described by Landis and Koch, 24 was employed. This rating system defines a score of less than 0 as “less than chance agreement,” 0.01 to 0.20 as “slight agreement,” 0.21 to 0.40 as “fair agreement,” 0.41 to 0.60 as “moderate agreement,” 0.61 to 0.80 as “substantial agreement,” and 0.81 to 0.99 as almost “perfect agreement.”
Results
Oxford Levels of Evidence
There was 100% agreement between the observers according to the Oxford Levels of Evidence (κ = undefined; P = undefined). Both observers 1 and 2 agreed that all 6 publications fulfilled the criteria to be considered level I.
Modified Coleman Methodology Score
The differences between the mean modified Coleman Methodology scores assigned by the 2 observers were not significant (45.0 compared with 45.3 points; P = .96). The 2 observers had agreement equivalent to chance regarding the modified Coleman Methodology scores, according to the Landis and Koch 24 categorical ratings (κ = 0; P = undefined), but did agree on 83% of studies.
When combining the assessments of both observers, the mean modified Coleman Methodology Score was 45.2 points (range, 34-67 points). Observer 1 assigned 5 (83%) of the studies a “poor” categorical rating and 1 (17%) of the studies a “fair” categorical rating. Observer 2 assigned all 6 (100%) of the studies a “poor” categorical rating. No study was assigned an “excellent” or “good” categorical rating. The modified Coleman Methodology Score breakdown for each reviewed study is listed below (Table 3).
Modified Coleman Methodology Score Breakdown by Reviewed Studies.
Revised CONSORT Statement
The differences between the 2 observers for the mean revised CONSORT Statement scores were not significant (13.7 compared with 14.3 points; P = .72). The 2 observers had a “fair” agreement regarding the revised CONSORT Statement scores according to the categorical rating (κ = 0.33; P = .27) but did agree on 67% of studies. 24
The mean of the revised CONSORT Statement scores for the 2 observers was 14.0 (range, 8-17 points). Observer 1 assigned 3 (50%) of the studies a “good” categorical rating and 3 (50%) of the studies a “fair” categorical rating. Observer 2 assigned 5 (83%) of the studies a “good” categorical rating and 1 (17%) of the studies a “fair” categorical rating. No study was assigned an “excellent” or “poor” categorical rating. The CONSORT score breakdown for each reviewed study is listed below (Table 4).
CONSORT Score Breakdown by Reviewed Studies.
Note. CONSORT = Consolidated Standards of Reporting Trials.
Discussion
Medical professionals often look to prospective randomized controlled trials to be gold standard scientific reports. Consequently, even astute readers might fall into the trap of thinking that simply because a study qualifies as being high-level evidence, it is inherently a high-quality study and thus has well-founded methods and results. This study was conducted to assess the quality of high-level prospective randomized control trials, specifically in the treatment of boxer’s fractures. We hypothesized that the current body of literature would be sound enough to support the conclusion that boxer’s fractures do not need reduction and can simply be treated with soft wrap techniques. However, similar to other studies analyzing the quality of evidence for scientific studies, 3 we found that the high-level evidence available was often unsatisfactory.
All 6 of the published studies that met our inclusion criteria were considered to be level I trials according to the Oxford Levels of Evidence. The modified Coleman Methodology Score further assessed the strengths and flaws of each study. Although these trials were considered high-level studies according to the Oxford Levels of Evidence, we found most studies were rated “poor” according to the modified Coleman Methodology Score, with only 2 (33%) being rated “fair” by 1 observer. The major limitations to the analyzed studies were numerous: Most failed to implement appropriate blinding, few had group comparability, most neglected to report clinical effect measurements, and only 1 reported the number of patients needed to treat. The only common strengths in these studies were reporting alpha error, having appropriate power, and not allowing cointerventions.
Our results regarding the revised CONSORT Statement were slightly better than the results for the modified Coleman Methodology Score. All the studies were rated as “good” or “fair,” with most of the studies rated “good” (50% and 83%) and the remainder rated “fair” (50% and 17%). Nearly all studies were able to earn full points for simple categories, such as having an appropriate title and abstract, background, inclusion criteria, and explanation of statistical methods. However, most neglected to state objectives or hypotheses. More importantly, almost every study failed to adequately conduct or properly explain randomization, generate appropriate allocation into treatment arms, or conduct suitable blinding. When critically assessed, these flaws in study design are oversights and can have harmful effects on the quality of study design.
The strengths of this study were the independent analyses of 2 qualified observers. Furthermore, we implemented 3 grading criteria that are quite elaborate and have been demonstrated in the past to adequately analyze scientific studies.3,19-22 Using the Modified Coleman Methodology Score and the revised CONSORT Statement, our reviewers were able to grade each study with statistically insignificant variability, according to Student t-test analysis using a quantitative score for each method. However, 1 weakness of our studies was the poor interobserver reliability according to the kappa statistic for categorical rating systems between scoring methods. This can most likely be explained due to the limited number (6) of high-level trials available for review regarding nonoperative boxer’s fracture management. Nevertheless, both observers did agree on 67% and 83% of studies for each method of analysis, which some authors argue is a better assessment of interobserver reliability than the kappa statistic when raters are well trained and little guessing is likely to exist. 25
Our data and analysis show that high-level studies do not always ensure that appropriate criteria are met, at least concerning the modified Coleman Methodology Score and revised CONSORT Statement. Regarding our hypothesis, we believed that current literature supporting nonoperative management with soft wrap techniques for boxer’s fractures would be high quality, but we found this information unsatisfactory when critically reviewed. This finding highlights a potential opportunity for future high-quality, high-level studies analyzing nonoperative management of boxer’s fractures. This review, similar to those published in the past, 3 serves as a further reminder to critically assess scientific studies that influence clinical decision-making, regardless of study design or level of evidence.
Footnotes
Ethical Approval
This study was approved by our institutional review board.
Statement of Human and Animal Rights
The making of this study did not require the use of patient informed consent, given the study design.
Statement of Informed Consent
This was not an experimental study involving humans or animals and thus did not need to undergo institutional review.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
