Abstract
Background:
The Functional Movement Screen (FMS) is utilized by professional and collegiate sports teams and the military for the prevention of musculoskeletal injuries.
Hypothesis:
The FMS demonstrates good interrater and intrarater reliability and validity and has predictive value for musculoskeletal injuries.
Study Design:
Systematic review and meta-analysis.
Methods:
A systematic review and meta-analysis were conducted using a computerized search of the electronic databases MEDLINE and ScienceDirect in adherence with the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines. Extracted relevant data from each included study were recorded on a standardized form. The Cochran Q statistic was utilized to evaluate study heterogeneity. Pooled quantitative synthesis was performed to measure the intraclass correlation coefficient (ICC) for interrater and intrarater reliability, along with 95% CIs, and odds ratios with 95% CIs for the injury predictive value for a score of ≤14.
Results:
Eleven studies for reliability, 5 studies for validity, and 9 studies for the injury predictive value were identified that met inclusion and exclusion criteria; of these, 6 studies for reliability and 9 studies for the injury predictive value were pooled for quantitative synthesis. The ICC for intrarater reliability was 0.81 (95% CI, 0.69-0.92) and for interrater reliability was 0.81 (95% CI, 0.70-0.92). The odds of sustaining an injury were 2.74 times with an FMS score of ≤14 (95% CI, 1.70-4.43). Studies for validity demonstrated flaws in both internal and external validity of the FMS.
Conclusion:
The FMS has excellent interrater and intrarater reliability. Participants with composite scores of ≤14 had a significantly higher likelihood of an injury compared with those with higher scores, demonstrating the injury predictive value of the test. Significant concerns remain regarding the validity of the FMS.
Professional and National Collegiate Athletic Association (NCAA) sports teams and the United States (US) military rely on physically healthy people to compose their respective work forces. Musculoskeletal injuries are a major source of lost participation time, lost income, and medical resources for the care of these injuries. 35 The Functional Movement Screen (FMS) is a screening test that was developed with the goal of identifying deficits in movements that may predispose an otherwise healthy person to injuries during activity.6-9 Preparticipation examinations have long been used to assess a person′s ability to safely participate in physical activity at the time of examination, but no existing screening test has been shown to predict a person′s risk of injury while participating in future activities. Kiesel et al 21 were the first to explore the possible predictive value of the FMS when they found that lower FMS scores were predictive of a significantly higher risk of injury in professional football players in a 2007 study. The value of such a screening test was quickly realized, and the FMS was widely adopted in organizations such as the National Football League (NFL), the National Hockey League (NHL), and the US military.20,21,28,33,35
Subsequent studies regarding the FMS, however, have produced varied results in regard to the injury predictive value as well as the validity of the FMS as a screening test.4,11,14,15,20,39,40 Analyses of the structure of the FMS have questioned the ambiguity inherent in its grading structure and its ability and sensitivity to identify functionally relevant movement limitations.1,2,5,14,15,40 Interrater and intrarater reliability were also introduced as possible sources for error in the FMS, although some early studies found high reliability among examiners with varying levels of experience. 27
Given the implementation of the FMS in numerous organizations and the growing body of literature examining the FMS, we performed a systematic review of the literature and meta-analysis to determine whether (1) the FMS is a reliable screening tool; (2) the FMS is a valid tool to identify functional asymmetries; and (3) if a lower score, and what specific score, on the FMS correlates with a higher risk of musculoskeletal injuries. We hypothesized that the FMS demonstrates both interrater and intrarater reliability with validity that can be used to identify people at a higher risk of an injury during activity.
Methods
A written protocol was developed in adherence to the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines to conduct a systematic review and meta-analysis of the available literature. 24 The MEDLINE and ScienceDirect electronic databases were searched for relevant studies with the primary search terms “Functional Movement Screen” OR “FMS” and secondary search terms “reliable” OR “reliability,” “valid” OR “validity,” “effective” OR “effectiveness,” “predict” OR “prediction,” “injury predict” OR “injury prediction,” and “predictive value” OR “injury predictive value.” Inclusion criteria included (1) English-language studies in peer-reviewed journals and (2) use of the FMS to assess uninjured people before participation in their respective activities. Any reviews, case reports, technique articles, or abstracts were excluded. The references of articles that met inclusion/exclusion criteria were also hand reviewed to ensure that any additional relevant studies were not missed.
Duplicates were removed from the results of each of the 3 separate searches. The titles and abstracts for all of the studies were then screened by the senior author (A.D.) to remove studies that did not involve the FMS. Each relevant study was then reviewed by the senior author and assessed using our inclusion and exclusion criteria for appropriateness for qualitative and quantitative analyses in our study. Qualitative and quantitative analyses were then performed by all authors. Quantitative analysis was performed for all articles in which data were sufficient to be included with other studies in our meta-analysis as described below. A diagram of our search methodology can be found in Figure 1.

Search methodology.
Data Extraction and Statistics
Interrater and Intrarater Reliability
Interrater and intrarater reliability were assessed using the intraclass correlation coefficient (ICC). Where variance was not directly reported, the confidence interval was used to determine the variance using the Fisher method.13,19 An evaluation for heterogeneity using the I2 statistic suggested significant heterogeneity between studies. A meta-analysis was performed using the DerSimonian and Laird 10 random-effects model. A sensitivity analysis was performed for the inclusion of the study of Smith et al, 37 which reported 4 separate analyses of intrarater reliability (1 for each subtype of observer). As Smith et al 37 included the individual ICC for each rater, the data were pooled to reduce the analysis to a single ICC, which was then included in the meta-analysis of the ICC.
Injury Predictive Value
The pooled effect measure that was subjected to the meta-analysis was the odds ratio (OR) for failure of functional movement. The numerical cutoff of the FMS was evaluated at a value of ≤14, which was the only value consistently utilized in all of the studies that met inclusion/exclusion criteria. Again, there was significant heterogeneity between studies based on the Mantel-Haenszel Q statistic. Therefore, the DerSimonian and Laird 10 random-effects model was used to estimate the pooled OR. 26 All analyses were performed using R (version 3.1.3) statistical software 32 and the rmeta package. 25
Validity
Because of insufficient reporting and heterogeneity of the data, a pooled quantitative analysis of validity could not be performed. Any reported data are included as part of the results and discussion of the articles from which they are extracted as part of our qualitative review.
Results
The initial search using the primary search terms resulted in 111 articles. Inclusion of the secondary search terms and removal of duplicates resulted in 45 articles addressing reliability, 43 articles addressing validity, and 33 addressing the injury predictive value, with several articles addressing more than 1 aspect of the FMS. Inclusion and exclusion criteria resulted in 11 studies evaluating reliability for qualitative analysis, which can be found in Table 1, and 6 studies for quantitative analysis/data synthesis.
Reliability Studies a
FMS, Functional Movement Screen; ICC, intraclass correlation coefficient; Inter, interrater; Intra, intrarater.
95% CI not reported.
Results that did not include the ICC with associated 95% CIs were unable to be included in the meta-analysis. Inclusion and exclusion criteria resulted in 9 studies evaluating the injury predictive value for qualitative analysis and 9 for quantitative analysis, which are described in Table 2. Inclusion and exclusion criteria resulted in 5 studies evaluating validity for qualitative analysis (Table 3). Given the variability of these studies for validity, a quantitative analysis could not be performed.
Injury Predictive Value Studies a
FMS, Functional Movement Screen; NCAA, National Collegiate Athletic Association; OR, odds ratio.
Validity Studies a
FMS, Functional Movement Screen; NCAA, National Collegiate Athletic Association.
Reliability
Interrater Reliability
Ten studies evaluated interrater reliability of the FMS. All studies examined the reliability of scores of more than 1 examiner grading the same participants. Significant differences were seen in the characteristics of the raters included in the studies. Five of the 10 studies included examiners of various levels of experience with the FMS. Only a few studies included raters specifically certified in the FMS.
Nine of the 10 studies found acceptable interrater reliability, with ICC values of 0.76 to 0.98. Shultz et al 36 was the only study to report poor interrater reliability with a Krippendorff alpha (α) value of only 0.38, which is well below the 0.8 cutoff considered acceptable. The Cohen kappa (κ) coefficient was reported as a measure of reliability for each individual FMS test component in 6 studies. These values varied, despite most finding overall scores had acceptable interrater reliability. Of the individual test components, the in-line lunge, rotary stability, and the hurdle step were all implicated as the least reliable component by at least 1 study.27,29,30,34,38,41
There were 5 studies that measured interrater reliability among raters with varying levels of FMS experience. Only 1 of those found unacceptable interrater reliability. Additionally, Gulgin and Hoogenboom 18 and Shultz et al 36 found that overall scores did not differ significantly between raters of different experience levels. Shultz et al 36 found overall unacceptable interrater reliability, but further analysis of their data did not show that experience had any affect. Both interrater reliability of raters with less than 1 year of experience (ICC, 0.44; 95% CI, 0.12 to 0.67) and that of raters with more than 2 years of experience (ICC, 0.177; 95% CI, –0.15 to 0.46) were unacceptable.
Five studies were finally included in the meta-analysis for interrater reliability as the remaining 5 studies did not have sufficient data for inclusion as previously described. Pooled quantitative analysis demonstrated that the mean ICC was found to be 0.81 (95% CI, 0.70-0.92) (Figure 2), indicating acceptable interrater reliability.

Analysis of interrater reliability. ICC, intraclass correlation coefficient.
Intrarater Reliability
Six studies evaluated intrarater reliability of the FMS. All studies examined the reliability of scores of the same examiners grading the same participants at 2 points separated by times ranging from 48 hours to 4 weeks. Five of 6 studies included multiple examiners of various levels of experience. Four of 6 studies had examiners evaluate video-recorded FMS tests, while 3 had examiners evaluate participants in real time.
All studies found acceptable intrarater reliability, with ICC values ranging from 0.6 to 0.96. As with interrater reliability, the level of experience did not consistently affect intrarater reliability. Gribble et al 17 showed that intrarater reliability increased with experience. However, Smith et al 37 found no difference in their 4 raters and actually found that the only certified FMS rater among their group had the lowest intrarater reliability. Individual components of the FMS showed significant variability with regard to intrarater reliability, again similar to the findings for interrater reliability.
A meta-analysis of intrarater reliability was performed, and a pooled mean ICC of 0.77 (95% CI, 0.58-0.96) (Figure 3) was obtained, which signified acceptable intrarater reliability. A synthesis of the data presented by Smith et al 37 (Figure 4) was included in the analysis, and an ICC of 0.81 (95% CI, 0.69-0.92) was observed, which also signified acceptable intrarater reliability (Figure 5).

Analysis of intrarater reliability excluding Smith et al. 37 ICC, intraclass correlation coefficient.

Analysis of intrarater reliability of raters per Smith et al. 37 ICC, intraclass correlation coefficient.

Analysis of total intrarater reliability. ICC, intraclass correlation coefficient.
Injury Predictive Value
Nine studies with a total of 2696 participants were identified that evaluated the injury predictive value of the FMS. All studies, per the inclusion criteria, evaluated healthy people before participation in their respective activities. As seen in the studies evaluating reliability, a significant variation was seen in the raters and in the populations being tested.
Definitions of injury were based on either time lost or treatment sought. The Kiesel et al,21,22 Butler et al, 3 and Dossa et al 11 studies used time lost from activity or sport, while Chorba et al, 4 Garrison et al, 16 Knapik et al, 23 O’Connor et al, 28 and Warren et al 39 based their definitions of injury on seeking medical care.
All 9 studies included in the quantitative synthesis compared participants who scored ≤14 with those who scored >14. This cutoff was first established by Kiesel et al 21 in a 2007 study in which a receiver operating characteristic (ROC) curve of their data showed that a cutoff of ≤14 maximized sensitivity and specificity and was affirmed in studies by Butler et al 3 and Garrison et al. 16 Chorba et al, 4 Dossa et al, 11 and Kiesel et al 22 used a cutoff of 14 based on the previous studies. O’Connor et al 28 and Warren et al 39 developed ROC curves with their data but found no value that optimized sensitivity and specificity and thus used 14 as a cutoff per prior studies as well. Knapik et al 23 reported data using 14 as a cutoff but found that sex affected optimal values. In their analysis of 1045 Coast Guard cadets, they found that the FMS score cutoff that maximized sensitivity and specificity, determined by the Youden index, was 11 for men and 14 for women. Additionally, the risk ratio was optimized at a cutoff of 12 for men and 15 for women.
Six of the 9 studies found that participants with an FMS score of ≤14 had a statistically significant higher risk of injury during subsequent activity than those with scores of >14. Studies consisted of mostly male participants, although Garrison et al 16 and Knapik et al 23 included female participants. The ORs ranged from 1.42 to 11.67. Three studies with a total of 225 participants did not find a statistically significant correlation between FMS scores and the risk of injuries. Chorba et al, 4 the only study with all female participants, as well as Dossa et al 11 and Warren et al, 39 found ORs that ranged from 1.01 to 3.85 but were not statistically significant.
A pooled quantitative synthesis using all 9 studies was performed using a score of 14 points as a cutoff cumulative score. An OR of 2.74 (95% CI, 1.70-4.43) was found (Figure 6), indicating that participants who scored ≤14 on the FMS had a 2.74 times greater probability of sustaining an injury during subsequent activity than those who scored >14 on the FMS.

Analysis of injury predictive value. OR, odds ratio.
Validity
Ten studies were identified in the initial search that evaluated the validity of the FMS. We defined validity based on the described purpose of the screen to identify deficiencies in movements that may contribute to a higher risk of injuries. 6 The application of the screen as an injury prediction tool is evaluated separately. The full text of all 10 studies was then reviewed independently by 2 of the authors (C.A.O., M.L.S.), who agreed independently on the appropriateness of 5 of the 10 studies for inclusion in the qualitative analysis. One study did not meet inclusion/exclusion criteria as it was an abstract only and was excluded. One study administered the FMS differently than it has been described in the literature and was excluded as well. Two studies were excluded as they sought to correlate scores with athletic performance, and the last excluded study looked at changes in FMS scores with training intervention, all outside our focus on the validity of the screen itself. A pooled quantitative synthesis could not be performed on these 5 studies because of the absence of a standard quantitative value, which is used to assess the validity of the FMS.
Kazman et al 20 used the FMS results of 934 Marine officer candidates to evaluate the factor structure of the FMS using the Cronbach alpha value and exploratory factor analysis (EFA). The Cronbach alpha value was found to be 0.39, less than 0.5, which signifies unacceptable internal consistency of the test. EFA additionally showed that the different test components within the FMS differ in their contributions to the total score, suggesting that the FMS is not unidimensional and cannot be used in its current form as a unitary construct to predict injuries.
Two articles evaluated the grading used in the FMS. Frost et al 14 hypothesized that instruction on grading criteria could improve the performer′s score on the FMS. Twenty-one firefighters completed the FMS first and then were educated on the specific grading criteria used to assess each move. After the instruction, the participants were tested again, and the score was found to be significantly higher for 4 of the 6 moves. Their conclusion was that additional performance instruction significantly changed FMS scores, not only functional asymmetries, speaking to potential flaws in the validity of the FMS. Whiteside et al 40 compared manual real-time testing to objective kinematic grading utilizing motion capture in the evaluation of 11 NCAA Division I athletes. They found significant differences between the 2 types of grading for all of the FMS exercises tested (only 6 of the 7 moves were tested), pointing to the ambiguous criteria for grading. This study also highlighted the difficulty in assessing multiple aspects of a movement from one vantage point.
Two studies evaluated the FMS against other measures of functional movement. Clifton et al 5 sought to validate the FMS by attempting to correlate scores with measurements in static balance during a single-leg stance: center of pressure velocity, center of pressure area, and center of pressure standard deviation in the medial-lateral and anterior-posterior directions. These measures had previously been validated to measure fatigue, which is a risk factor for musculoskeletal injuries. 5 They tested active people, defined as participants aged 18 to 50 years who exercised 3 times per week for at least 30 minutes per workout, both before and after exercise. They hypothesized that fatigue would lead to a decrease in static balance after exercise and would correlate with changes in FMS scores. Although the static balance measurements decreased after exercise as expected, they found that FMS scores did not change. They concluded that the FMS was not a useful predictor of who will experience greater balance deficits after exercise. This finding does not support the idea that the FMS can be used in all settings as an injury predictor.
Beach et al 1 hypothesized that if the FMS identified deficiencies in functional movements, scores may correlate with more activity-specific parameters. A total of 30 firefighters were evaluated: 15 who scored >14 on the FMS and 15 who scored ≤14. The participants were height and weight matched. They were asked to perform the standardized task of lifting a box. Lumbar spine loading magnitudes and lumbar spine angles were measured and compared with FMS scores. They did not find a statistically significant difference between the 2 groups, which calls into question the ability of the FMS to measure core stability.
Discussion
The popularity and utilization of the FMS have grown rapidly since its development, bolstered by evidence in the literature to suggest its injury predictive value. Its adoption at the highest level of athletics as well as the military and other public service organizations has further contributed to its rise in popularity, despite the existence of conflicting literature evaluating not only the injury predictive value but also the validity and reliability of the FMS. Given this, we sought to assess and assimilate the relevant literature where appropriate to determine whether the FMS is a reliable, valid screening test with an injury predictive value. It is essential that limitations in the screen are understood to eliminate to the extent possible inaccurate evaluations that can result in significant consequences for both the screened participant and his or her respective organization.
The reliability of the examiners has been thoroughly investigated, especially given the theoretical concern that varying levels of experience as well as the presence or absence in certification would result in significant differences in examinations. The variety in examiners evaluating participants undergoing the FMS and the methodology for how participants were examined are clearly evident in our included studies for both reliability and the injury predictive value, introducing numerous biases that could affect the results of the respective individual studies and, ultimately, the results of the meta-analysis. Overall, there is significant evidence that the composite FMS test is reliable and can be replicated by raters with varying degrees of experience with the FMS. Gribble et al 17 had the only study evaluating interrater or intrarater reliability that found that increasing experience led to increased reliability. Smith et al 37 was the only other study to report the ICC for interrater reliability of multiple raters who were both certified and not certified with different experience levels, and they found no difference. None of the 5 studies evaluating interrater reliability with raters of different experience levels found any effect on reliability. Only Shultz et al 36 found poor interrater reliability, but as stated above, further analysis found that dividing more and less experienced raters did not improve their interrater reliability. The results of our meta-analysis, which showed high interrater and intrarater reliability, suggest that level of experience and formal certification by Functional Movement Screen Inc have little effect on scoring of the FMS.
Interrater and intrarater reliability for individual tests of the FMS did vary greatly among all types of raters, which may indicate a lack of specificity in the grading criteria or simply a level of difficulty in grading certain subtests that could be a source of error in the screen. This may be an area for further research. While the ICC is used as a quantitative measure of reliability, individual test subsection analysis was often reported using the kappa coefficient. While it would have been ideal to analyze each individual subtest within the FMS, a degree of variance was often not reported with the kappa coefficient. Because of these restraints in statistical methodology, we were unable to calculate a P value for our calculations based on insufficient reporting of raw data and/or variance.
Our study demonstrates that the historical pass/fail cutoff of 14 points is valid in predicting those at a higher risk of injury. Evidence supporting 14 points as the optimal cutoff, however, is limited as only 2 studies replicated the finding of Kiesel et al 21 via independent ROC curves. These findings were also limited to only male participants, which may be a limitation. Knapik et al 23 found that the optimal cutoff may differ by sex. O’Connor et al 28 and Warren et al 39 independently found no optimal value in their data. Studies with mixed-sex populations did find statistically significant results, but the effect of sex and other population characteristics on a cutoff that optimizes the injury predictive value of the FMS may be an area of further research.
Garrison et al 16 found that a history of injuries alone could identify those at a higher risk of future injuries and that combining a history of injuries with an FMS score of ≤14 suggested a significantly higher risk of future injuries, increasing their ORs from 5.61 (95% CI, 2.73-11.51) to 15.11 (95% CI, 6.60-34.61). This is consistent with the prior finding that a history of injuries is associated with lower FMS scores. 31 Further research may identify other factors that, combined with the FMS, significantly affect its injury predictive value.
The included studies regarding the validity of the FMS point to several concerns about its structure and its ability to detect abnormal movement patterns. Kazman et al 20 showed that the composite score of the FMS is not valid as a unitary construct as it is often used. Grading may be flawed by somewhat ambiguous criteria, and Frost et al 14 showed that educating those being screened on the criteria can significantly affect scores, suggesting that scores may be reflecting more than just the physical characteristics that they intend to assess (ie, learned behavior). Comparisons with other measures of movement and balance also did not find a correlation to FMS scores, again questioning its accuracy or at least sensitivity in detecting physical abnormalities. Because of the absence of any gold-standard comparison and significant heterogeneity of the existing data, it is difficult to derive any definitive conclusions from the current literature as to whether the FMS is a valid tool for the measurement of functional limitations and asymmetries.
Limitations of our meta-analysis of both reliability and the injury predictive value included insufficient reporting of raw data, P values, and variance. Other observed potential biases in analysis include the need to determine variance using the Fisher method in situations where these data were not directly reported. 13 This has been well described previously in the literature. 13
Although all studies included either time missed from sport or activity because of an injury or that which required medical attention, the variability in the definition adds an element of heterogeneity to our analysis. As these studies include a variety of participants including athletes, military personnel, and public servants such as firefighters, much of this variability is inherent to their varying respective activities and demands.
Based on the results of this systematic review and meta-analysis, the FMS as a composite score has excellent interrater and intrarater reliability and can be effectively administered by raters of varying levels of experience with the FMS both with and without formal certification. Participants who score ≤14 on the FMS have greater than twice the odds of sustaining a musculoskeletal injury than those with scores of >14. However, the FMS lacks validation of its structure as a composite score of multiple subtest scores and of its ability to accurately and sensitively measure deficits in posture and balance. Despite this demonstrated injury predictive value of the FMS, the clinical application of the FMS should be exercised with caution until further studies can confirm the screen′s validity.
Footnotes
One or more of the authors has declared the following potential conflict of interest or source of funding: A.D. is a consultant for Smith & Nephew and is on the speakers’ bureau for Smith & Nephew and Biomet.
An online CME course associated with this article is available for 1 AMA PRA Category 1 Credit™ at
. In accordance with the standards of the Accreditation Council for Continuing Medical Education (ACCME), it is the policy of The American Orthopaedic Society for Sports Medicine that authors, editors, and planners disclose to the learners all financial relationships during the past 12 months with any commercial interest (A ‘commercial interest’ is any entity producing, marketing, re-selling, or distributing health care goods or services consumed by, or used on, patients). Any and all disclosures are provided in the online journal CME area which is provided to all participants before they actually take the CME activity. In accordance with AOSSM policy, authors, editors, and planners’ participation in this educational activity will be predicated upon timely submission and review of AOSSM disclosure. Noncompliance will result in an author/editor or planner to be stricken from participating in this CME activity.
