Abstract
Variance inflation factors (VIF scores) are regression diagnostics commonly invoked throughout the social sciences. Researchers typically take the perspective that VIF scores below a numerical rule-of-thumb threshold act as a “silver bullet” to dismiss any and all multicollinearity concerns. Yet, no valid logical basis exists for using VIF thresholds to reject the possibility of multicollinearity-induced type 1 errors. Reporting VIF scores below a threshold does not in any way add to the credibility of statistically significant results among correlated variables. In contrast to this “threshold perspective,” our analysis expands the scope of a perspective that has considered multicollinearity and misspecification. We demonstrate analytically that a regression omitting a relevant variable correlated with included variables that exhibit multicollinearity is susceptible to endogeneity-induced bias inflation and beta polarization, leading to the possible co-existence of type 1 errors and low VIF scores. Further, omitting variables explicitly reduces VIF scores. We conclude that the threshold perspective not only lacks any logical basis but also is fundamentally misleading as a rule-of-thumb. Instrumental variables represent one clear remedy for endogeneity-induced bias inflation. If exogenous instruments are unavailable, we encourage researchers to test only straightforward, unambiguous theory when using variables that exhibit multicollinearity, and to ensure that correlated co-variates exhibit the expected signs.
Keywords
Introduction
Multicollinearity, or more precisely, near-multicollinearity, occurs when two or more independent variables in a regression exhibit substantial zero-order bivariate correlations with each other (Wooldridge, 2016). Multicollinearity among regression coefficients remains an active topic of concern across the physical and social sciences after a half-century of debate: More than 473,000 Google Scholar papers mention the term “multicollinearity.” Variance inflation factors (VIF scores) are the most frequently employed diagnostics of multicollinearity in academic research, with 175,000 (37%) of the 473,000 multicollinearity papers on Google Scholar mentioning their use. In the year 2010, 4,840 (28%) of 17,300 papers that mentioned multicollinearity also referenced VIF scores. In contrast, 32,900 (68%) of all 48,600 papers that mentioned multicollinearity in 2021 also referenced VIF scores. In 2022, 34,500 (78%) of 44,100 referenced VIF scores. 1 We conclude from these numbers that VIF scores have become near-mandatory diagnostics that researchers feel compelled to present whenever multicollinearity concerns arise.
Methodologists have long known that multicollinearity is associated with imprecise estimation. The ubiquitous and technically correct but nonetheless highly misleading statement that “multicollinearity does not bias coefficients” originates from this perspective. Variances of β coefficients become inflated along with VIF scores. The result of large VIF scores may be type 2 errors (e.g., Belsley et al., 1980; Goldberger, 1991). The fundamental problem in practice is that, as we will demonstrate, no valid logical basis exists for using VIF thresholds to reject the possibility of multicollinearity-induced type 1 errors. When researchers report low VIF scores in practice they are almost universally doing so to buttress the credibility of statistically significant results. Implicitly, they are trying to claim that low VIF scores reduce the possibility of type 1 errors. A potentially valid diagnostic for type 2 errors has been misappropriated broadly across the social sciences as a “silver bullet” to dismiss the possibility of type 1 error.
More recent work, however, has in fact focused explicitly or implicitly on possible type 1 error. In particular, the previously ignored role of multicollinearity not in terms of creating bias outright but, rather, of inflating existing biases has been examined (Kalnins, 2018, 2022; Middleton et al., 2016; Steiner & Kim, 2016; Winship & Western, 2016). The analyses of Kalnins (2018, 2022, 2024) analyze not only coefficient bias but also multicollinearity effects on standard errors (SEs) and t-statistics. These additional derivations are prerequisites for an assessment of the likelihood of excess type 1 errors.
Further, Kalnins (2018) documented the phenomenon of “beta polarization” among published papers in empirical organizational research where estimated β coefficients of positively correlated variable pairs are large in absolute magnitude, statistically significant, and opposite in sign. One of the coefficients often exhibits a counter-intuitive sign. Specifically, Kalnins (2018) identified 64 articles published in the Strategic Management Journal (SMJ) between 2013 and 2017 that contained at least one result reported as providing meaningful support for a hypothesis, but that may be a type 1 error based on the beta polarization criteria. In 41 of the 64 articles, at least one hypothesized coefficient affected by multicollinearity flipped sign from that of its bivariate correlation with the dependent variable. Kalnins (2018) proposed the “common factor” data-generating process (DGP) as a source of both beta polarization and excess type 1 errors in multivariate regression results.
In this paper, we extend Kalnins’ (2018) results regarding bias and excess type 1 errors to the general case of any regression analysis where endogeneity is present in the form of an omitted variable. No precise assumptions regarding DGP are required. We contribute to the social sciences research methods knowledge base by demonstrating how multicollinearity can inflate endogeneity-induced biases, polarize beta coefficients, and create excess type 1 errors, all in the presence of low VIF scores. This is an important contribution beyond Kalnins (2018) because, as Lindner et al. (2022) point out, it is difficult to determine the likelihood that a specific DGP has created a dependent variable.
We have organized this paper as follows. We first describe the three perspectives regarding the use of VIF scores as multicollinearity diagnostics in empirical research: The threshold perspective, the t-statistic perspective, and the bias inflation perspective. Second, we present a methodological practice review of how organizational researchers use VIF scores in practice. The possible role of multicollinearity in causing bias inflation rather than variance inflation is the overwhelming if largely implicit concern in published work. Third, we examine the standard formulas for the VIF score and for omitted variable bias (OVB) in multivariate regression to show that multicollinearity may lead to near-infinite biases and polarized β coefficients. We then use numerical examples of moderate multicollinearity and small-to-medium effect sizes of independent variables to generate conclusions about OVBs inflated by multicollinearity. We conclude that empirical researchers should not use VIF scores as multicollinearity diagnostics. Unfortunately, there is no obvious statistic that we could call the “BIF score” to detect bias inflation. Instead, researchers should rely on bivariate, zero-order correlations and the appearance of counter-intuitive findings to detect possible multicollinearity concerns. We then recommend that researchers mitigate these concerns either by seeking out an exogenous instrumental variable, or by providing clear, straightforward, unambiguous theory for variables of theoretical interest, and by establishing consistency of correlated co-variate coefficients with previous results and theory.
The Three Perspectives on VIFs
Typically, what researchers claim based on a presentation of VIF scores is that, if all of a regression's VIF scores are below a threshold such as 5 or 10, they may dismiss multicollinearity concerns regarding statistically significant variables without further consideration. In other words, the VIF scores are a supplementary analysis used to support a conclusion that statistically significant variables are not type 1 errors. We refer to such conclusions based on VIF scores being below thresholds as the “threshold perspective.”
Methodologists have mixed views regarding the efficacy of VIF thresholds. On the one hand, Hair et al. (2006, p. 230) recommend using thresholds and specifically a threshold of ten, though they suggest that “when sample sizes are smaller, the researcher may wish to be more restrictive.” On the other hand, Wooldridge (2016, p. 86) states “setting a cutoff value for VIF above which we conclude multicollinearity is a ‘problem’ is arbitrary and not especially helpful.” We will demonstrate that the arbitrariness of VIF thresholds is far from their greatest flaw. Large VIF scores may indeed suggest type 2 errors when coefficients do not exhibit statistical significance, in a logically similar sense to small sample sizes (Goldberger, 1991). But neither small nor large VIF scores provide information about the possible presence of type 1 errors.
The second perspective comes from analyses conducted by O’Brien (2007) and Lindner et al. (2020). These papers argue that researchers should not use VIF rule-of-thumb thresholds either to dismiss (below threshold) or suggest (above threshold) the possibility of multicollinearity problems. T-statistics are robust to the presence of multicollinearity in their view, and therefore supplementary diagnostics such as VIF scores are superfluous. Lindner et al. (2020, p. 296) state “In a regression model, VIFs represent the degree to which regression coefficient variance is too large, i.e., larger than in a model with only independent variables with zero partial correlation. Consequently, significant results under high variance inflation may be taken to be conservative.” That is, under this perspective, statistically significant coefficients indicate that multicollinearity has not affected results and thus is not an issue. We refer to this logic as the “t-statistic perspective.”
Kalnins (2018, 2022) introduced a third perspective regarding VIF scores using dependent variables generated by the Common Factor DGP. This DGP may be commonplace in reality but is inconsistent with the Gauss–Markov assumptions. Kalnins (2018) presented examples where variable coefficients with large t-statistics, even when accompanied by low VIF scores, were in fact type 1 errors. Multicollinearity may in this case also cause beta polarization, that is, it may artificially push β coefficients of positively correlated variables in opposite directions. If the correlations are negative, it may push β coefficients in the same direction. VIF scores may remain low despite these biases. We refer to this as the “bias inflation” perspective.
Methodological Practice Review of Published Organizational Research
To determine the importance of multicollinearity and VIF scores in practice in organizational research, we examined 370 articles published in 2020 or 2021 in the following journals: Academy of Management Journal, Journal of Applied Psychology, Journal of Management, and SMJ. For each paper, we recorded three matters of interest. First, we checked whether any hypothesis-testing variables had bivariate zero-order correlations with other variables greater than 0.30. While 0.30 is by no means a minimum threshold for problematically high correlations, it is fully capable of producing excess type 1 errors in realistic scenarios (Kalnins, 2018). Second, we checked whether the authors considered potential multicollinearity problems. Finally, we noted whether these authors used VIF scores in their analysis.
Of the 370 papers reviewed, 200 had at least one hypothesized variable with a bivariate correlation greater than 0.30 with another variable in the regression. Of the 200 problematic papers, 118 either made no mention of multicollinearity or made no attempt to determine whether multicollinearity may have influenced results. However, 103 papers acknowledged that multicollinearity may have been an issue (not all are among the 200 with problematic correlations). Of the 103, 75 (73%) used VIF scores to evaluate concerns. This percentage equals exactly the average of the 2021–2022 percentages that mention VIF scores among the 92,000 + papers on Google Scholar across all disciplines that reference multicollinearity.
In all but one of our 75 organizational papers that employed VIF scores, authors used what they considered to be low scores (below thresholds of 5, 10, or even 20) to summarily dismiss multicollinearity concerns regarding statistically significant variables of theoretical interest. Implicitly or explicitly, then, authors used low VIF scores to buttress confidence that statistically significant results were legitimate and not type 1 errors. The concern here is overwhelmingly one of bias inflation rather than variance inflation.
Fifty-two papers explicitly stated thresholds such as 5, 10, and even 20. This analysis shows that, first, as suggested by the considerable number of papers with high correlations, multicollinearity represents a potential problem in organizational research that often goes unacknowledged. Second, for those that do consider multicollinearity in their analysis, authors almost exclusively use VIF scores below thresholds to validate support for hypotheses and statistically significant results as per the threshold perspective.
VIFs and Omitted Variables
VIFs are one component of the standard formula for the estimated variance of β coefficients, with which they have a linear and positive relationship. There is a separate VIF score for each independent variable in a regression. VIF scores are a function of independent variables included in a regression; they have no relationship with the dependent variable. We can describe VIF scores in two ways, both equally correct. First, the VIF score associated with each independent variable xi is the diagonal element ii of the inverse correlation matrix of all independent variables (Marquardt, 1970). Second, the VIF score associated with xi is the sum of squared values of xi divided by the sum of squared residuals of xi that results from a supplementary regression of xi on all other independent variables x∼i (Smith & Campbell, 1980):
VIF scores function as multipliers of the variances of the β coefficients in a regression where independent variables are correlated, relative to a hypothetical case where those same variables are uncorrelated. The minimum value of a VIF score is one when all independent variables in a regression are uncorrelated. This is considered ideal; a VIF score can never be “too small.” The only possible implication for a set of results, then, is whether the VIF scores are “too large.” Large VIF scores are indicative of a large variance of the β coefficients.
This basic background discussion provides sufficient information to deduce our first conclusion. The equation for the VIF score and the interpretation of Smith and Campbell (1980) makes it clear that the less the independent variables explain in total about each other, the lower the Ri2, and thus the lower the VIF score. If the Gauss–Markov assumptions are fulfilled, implying no omitted variables, low VIF scores may indeed suggest a regression with stable β coefficients. However, assuming that variables are omitted, low VIF scores may paradoxically represent greater misspecification in the form of greater endogeneity and more omitted variables.
Analysis of OVB and VIF Scores
It is well known that omitted variables bias regression coefficients. Most texts focus on the simple case of one omitted variable and only one included variable. In this case, the bias will remain finite, yielding a β coefficient that is an additive combination of the true effects of the included and omitted variable. Methodological texts such as Wooldridge (2016) and Maddala (2001) present the more general form of bias from an omitted variable on multiple included independent variables. Those authors caution that biases on different included variables might go in different directions.
The multivariate form of OVB is a violation of the Gauss-Markov assumptions because variable(s) present in the DGP of the dependent variable y are absent from the regression analysis. In the most general form of OVB (see, e.g., Angrist & Pischke, 2009), let
The term
Our Model: The Role of the VIF Score in OVB
We assume there are two standardized variables x1 and x2 in the “short” regression matrix X, which omits an important third variable, also standardized, and designated by Z. For even more simplicity, but with no loss of generality due to the independence of the bias from the observable variables’ true effects on dependent variable y, we will assume that those true effects are
The observable variables in X can be viewed as proxy variables for Z; in this simplest case they have no direct effects of their own on DV y. However, all our results generalize to the case where the variables in X have their own distinct but finite effects on the DV. Given our standardization assumptions,
The expected value of the estimated coefficients can be written as:

An illustration of equation 3 with corr(x1,Z) = 0.2 and corr(x2,Z) = 0.
If corr(x1, Z) ≠ corr(x2, Z), true in any empirical setting, the difference in the numerator remains finite while VIF approaches ∞. The bias approaches ±∞. Further, if θ → 1 then there will always exist a θ close enough to 1 such that the quantities (corr(x1,Z) – θ corr(x2,Z)) and (corr(x2,Z) – θ corr(x1,Z)) will necessarily have opposite signs. This combination yields
If θ → −1 there will always be a value of θ close enough to −1 such that the quantities (corr(x1,Z) – θ corr(x2,Z)) and (corr(x2,Z) – θ corr(x1,Z)) will necessarily have the same sign. This combination is an example of beta homogenization, again, for a more generalized case than research has previously demonstrated. The coefficients of both x variables will approach +∞ if their average correlation with Z is positive but will approach –∞ if negative.
We recognize that, as θ → 1, corr(x1,Z) and corr(x2,Z) must necessarily become more and more similar in magnitude. Otherwise, the correlations become infeasible because the correlation matrix of x1, x2, and Z will cease to be positive definite (Spanos & McGuirk, 2002). For a given fixed pair corr(x1,Z) ≠ corr(x2,Z), there will be a feasible maximum θ < 1 or minimum θ > –1 that is associated with a maximum feasible VIF.
VIF Scores and the t-Statistic
Before we can conclude that VIF scores may be associated with type 1 errors, not just inflated β coefficients, we must consider the SEs and the t-statistics (t) in the presence of an omitted variable. On the one hand, Kalnins (2022) demonstrated that, if the Gauss-Markov assumptions hold, then t → 0 whenever θ → 1. In this case, VIF → ∞, type 2 error becomes a certainty and type 1 error an impossibility. On the other hand, in the presence of an omitted variable, we will now show that the t-statistic may maintain a non-zero value when θ → 1 and may remain of an appropriate size to suggest statistical significance even when there is no true effect. We write the t-statistic for each of the two variables xi within X.
Under What Conditions Does OVB Generate a Type 1 Error?
From Equation 9, we observe that it is unclear whether VIF increases the likelihood of statistical significance and of type 1 errors, because the VIF term appears within square roots in both the numerator and the denominator. Further complicating matters, the bivariate correlation θ of the included independent variables, of which VIF is a function, also appears in both the numerator and denominator. This complexity suggests that there is no intuitive reason to associate high VIF scores with type 1 errors as the common practitioner usage of the thresholds implicitly suggests.
To gain more insight we investigate an illustrative special case. We consider the case where corr(x1,Z) ≠ 0 and corr(x2,Z) = 0, and where, again, the true effects on dependent variable y are
Numerical Examples with Varying Sample Sizes
We now apply Equations 10 and 11 using numerical examples. We choose correlations between the observable and omitted variables based on Cohen's (1988) “small” and “medium” effect sizes. Cohen (1988) proposed that correlations of 0.2, and, relatedly, R2 values of 0.04, constitute “small effects,” while correlations of 0.5, and R2 values of 0.25 constitute “medium” values. We present two tables below to show the sample sizes and correlations, along with the corresponding VIFs, that will create large expected values of t-statistics and thus type 1 errors for the coefficient of variable x2 in the estimated regression y = βX + e. Higher expected values of the t-statistic will result in a greater likelihood of type 1 errors. For example, when the expected value of the t-statistic = ±1.96 and the cutoff for significance is p < .05, the likelihood of a type 1 error is 50% because half of samples will have an absolute value of the t-statistic that is greater than this number.
t-Statistics for Estimated β Coefficients of x2; Corr(x2,Z) = 0. Small effects: corr(x1,Z) = 0.2, σ2 = 0.3 yielding Var(y) = 1.3 and R2 ≈ 0.04 at θ = 0.5.
Cells with absolute t-values greater than 1.96 represent type 1 errors with a greater than 50% likelihood.
t-Statistics for Estimated β Coefficients of x2; corr(x2,Z) = 0. Medium effects: corr(x1,Z) = 0.5, σ2 = 0.3 yielding Var(y) = 1.3 and R2 ≈ 0.25 at θ = 0.5.
Cells with absolute t-values greater than 1.96 represent type 1 errors with a greater than 50% likelihood.
In Table 1, corr(x1,Z) is set to 0.2, which will generate a “small” effect in a regression of y on X as per Cohen's classification, given that y = Z + e. We set corr(x2,Z) = 0. These are the same correlations depicted in Figure 1. Table 1 shows that even this small effect will cause type 1 errors for realistic sample sizes and correlations θ. In the first column, θ = 0.10 yields a VIF of 1.01, almost the absolute minimum value. The average estimated coefficient
Table 2, with medium effects, shows that type 1 errors become common at lower correlations and sample sizes and, not surprisingly, create larger β coefficients. In column 1 we observe that θ = 0.10 and the corresponding VIF of 1.01 yields an average t-statistic of ±1.96 when sample size N = 1,600, one-eighth of the 12,000 required for small effects. The average estimated coefficient
For those research outlets that prefer the presentation of results without relying on cutoffs for “significance,” such as p < .05, we emphasize that any inflation in the expected value of the t-statistic dramatically increases the probability of presenting results that appear rigorous but are dangerous distortions of the true measures.
t-Statistics for Estimated β Coefficients of x2; corr(x2,Z) = 0. N = 500.
Cells with absolute t-values greater than 1.96 represent type 1 errors with a greater than 50% likelihood.
Numerical Examples with Varying Differences Between Correlations
Table 3 shows the increasing likelihood of type 1 error that results from greater differences between the correlations of the independent variables with omitted variable Z. As in Tables 1 and 2, the average estimated t-statistic values come from Equation 11. We have fixed N = 500.
The second row of Table 3 demonstrates that a difference of 0.1 between corr(x1,Z) and corr(x2,Z) is insufficient to create excess type 1 errors until about θ = 0.70. The third row is the same as the fifth row of Table 1. The correlations with Z of 0.2 and 0.0 were those we used for Cohen's (1988) small effect. As we noted above, θ = 0.50 is the correlation between x1 and x2 for which we must seriously worry about type 1 errors for sample sizes in the 400–600 range when trying to detect small effects. The sixth row of Table 3 is the same as the third row in Table 2, where we used the 0.5 and 0.0 correlations as a basis for a medium effect. We observe high levels of type 1 errors for θ ≥ 0.30 when the difference between corr(x1,Z) and corr(x2,Z) is > 0.3, and for θ ≥ 0.10 when the correlation difference is > 0.7. While we cannot observe corr(x1,Z) and corr(x2,Z) in actual data due to the unobservable nature of Z, the occurrence of such large correlation differences in cases where θ ≥ 0.30 or even θ ≥ 0.10 may not be frequent. In general, we would expect that the larger the difference between corr(x1,Z) and corr(x2,Z), the smaller the correlation θ. We note that in the final row of Table 3, the correlation difference = 0.9 cannot be associated with any θ ≥ 0.43 because the full correlation matrix ceases to be positive definite.
An Alternative Mechanism That Links the VIF and Type 2 Error
We now return to the general t-statistic equation, Equation 9, to demonstrate the possibility not only of type 1 errors but also type 2 errors, based on the size of correlation θ. Consider a case with two effects: a medium and a small effect. Variable x1 is correlated with Z by an amount (0.4) twice that of the correlation between Z and x2. Importantly, both elements of X have positive correlations with the omitted Z. In this scenario, we can think of X as proxy variables for Z. Both should have positive coefficients, based on their positive correlation with Z.
Table 4 shows that small correlations θ such as 0.1 or 0.3 are associated with largely accurate levels of statistical significance, given a sufficient sample size: corr(x2,Z) is positive, and Z has a positive effect on y. Through the positive value of
t-Statistics for Estimated β Coefficients of x2; corr(x2,Z) = 0.2. Small effects: Corr(x1,Z) = 0.4, σ2 = 0.3 yielding Var(y) = 1.3 and R2 ≈ 0.12 at θ = 0.5.
What to do? Use Instrumental Variables
What, then, should be done, given that we have shown that any claimed efficacy of VIF scores as multicollinearity diagnostics is a myth? As we stated earlier, there is no obvious statistic that we could call the “BIF score” to detect bias inflation. In the absence of such a “silver bullet” we endorse the advice in Kalnins (2018, 2022): when bivariate correlations of a hypothesized variable with a covariate in the regression lie above an absolute value of 0.3, or above 0.2 for data sets with thousands of observations (see Tables 1 and 2), researchers should be concerned about multicollinearity. The use of zero-order bivariate correlations to detect multicollinearity has a drawback in that they cannot accommodate multicollinearity across several variables. However, the most deleterious effects of multicollinearity occur due to highly correlated variable pairs, as shown in this paper, and the bivariate correlation θ is the fundamental driver of such effects. Further, the bivariate correlation avoids the paradoxical relationship inherently built into diagnostics such as VIF scores of a low, seemingly acceptable value being positively related to a high likelihood of type 1 error.
Further, if one of the correlated variables exhibits a counter-intuitive effect, a sign of beta polarization, the researcher should be concerned about the bias-inflating role of multicollinearity. We recommend that researchers then check whether the counter-intuitive regression coefficient has flipped its sign from its bivariate correlation with the dependent variable. If yes, the study should precisely and transparently identify those additional co-variates that cause the flip when included (Lenz & Sahn, 2021). Can the pre-flip and post-flip signs of all relevant co-variates be clearly explained using existing theory and previously published results? Should concerns remain, we consider two approaches to avoid the possibility of publishing results that may well be type 1 errors. First, because we have demonstrated that the inflation of bias results from the violation of strict endogeneity in the form of omitted variables, we can apply the solution of instrumental variables (IV). Nobel laureate Joshua Angrist and co-author Jorn-Steffen Pischke (2009, p. 115) stated that “Undoubtedly, however, the most important contemporary use of IV methods is to solve the problem of OVB.” Therefore, if possible, we recommend finding one or more instruments for a theorized variable when either this variable or a correlated covariate appears counter-intuitive, and their correlation is above 0.30. Truly exogenous instruments will eliminate the bias from the omitted variable as well as the bias inflation caused by the multicollinearity. However, we note that the instrument in this case must be uncorrelated with, and independent from, not only the omitted variable but also from the correlated co-variate. Even if the theorized variable is completely uncorrelated with the omitted variable, an instrument is still required if the correlated co-variate is in fact endogenous through the omitted variable. If both correlated variables are theorized variables of interest, and at least one appears endogenous, then researchers will need to instrument both variables.
We appreciate that valid instruments are hard to find. Semadeni et al. (2014) conducted simulations and found two important results. First, instruments only weakly correlated with the endogenous variable result in high SEs and a high likelihood of type 2 errors. Second, inappropriate instruments that are not truly exogenous may result in statistically significant coefficient estimates that differ from their true values and thus could be type 1 errors. If not executed perfectly, the “cure” of instrumental variables may be no better than the “disease” of endogeneity.
What to do if Valid Instruments Cannot be Found?
A major implication of the conclusions presented throughout this paper is that data sets where substantial bivariate zero-order correlations exist (> |0.30|) between variables of theoretical interest and co-variates have limited potential for testing theory. For large data sets, such as those with N > 10,000, we may even be concerned with bivariate zero-order correlations > |0.20|. In cases where such correlations exist and where an appropriate instrument cannot be found, authors should take extra care to present only clearly falsifiable, unambiguous, straightforward theory-based hypotheses. Data with multicollinearity should not be used to test theory with counter-intuitive twists. Counter-intuitive theory-supporting results may well be nothing more than type 1 errors. Similarly, “horse-race” tests of opposing hypotheses are likely to result in type 1 error being the winning horse. HARKing (hypothesizing after results are known) is particularly dangerous when multicollinearity is present, because type 1 errors may “replicate” if future authors include the same correlated co-variates in future studies. What a field of inquiry may deem to be a robustly tested phenomenon may merely be a series of type 1 errors due to similar correlational structures across multiple data sets.
In addition to the greater-than-usual care that is required when hypothesizing, the presence of high bivariate correlations also behooves us to consider and investigate the results of co-variates. Authors should present clear evidence that the correlated covariate's coefficient (when focal variable is included) is not only consistent in sign with existing theory but also consistent in sign and magnitude with existing results from other sources. Many papers ignore co-variate results that make no sense. Editors and reviewers should demand a full supporting explanation for the signs and magnitudes of correlated co-variates, and whether their coefficients flip sign from their zero-order bivariate correlations with the dependent variable.
Kalnins (2018) recommended that authors present additional specifications of results to provide reviewers and readers with definitive confidence that statistically significant results are robust in the face of multicollinearity that stems from a Common Factor DGP. While his approach is not conclusive for our more general case of omitted variables, the presentation of multiple specifications remains of value. As per Conclusion 2, a positive correlation of two correlated variables may result in artificial “beta polarization” of their coefficients. One coefficient will become larger but remain positive when the second variable is included. The other variable will be smaller and may flip the sign when the first variable is included. If a paper's authors wish reviewers and readers to accept the support of a hypothesis, they must also convince us that the coefficient of the co-variate has the correct sign and a reasonable magnitude. Otherwise, we should not assume that the result including the co-variate is preferable, even though that would be true if the regression fulfilled the Gauss–Markov assumptions.
Vatcheva et al. (2016) presented an example that illustrates this dilemma using real-world medical data with blood pressure as the dependent variable. The addition of a correlated hypothesized variable, “body mass index,” flips the “waist circumference” coefficient from statistically significant in a theoretically supported, positive direction to marginally significant (p = .0539) in the negative, counter-intuitive direction. In the medical field, the research community sufficiently understands blood pressure such that they can definitively dismiss this “counter-intuitive” finding as nonsense. Yet, if a research community insists on the application of the threshold perspective, we would have to conclude that multicollinearity does not pose any concern because the VIF scores are below the threshold of five. We would then have to accept as valid the nonsensical result that waist circumference is negatively related to blood pressure. Similarly, if the research community applied the t-statistic perspective, we would again have to accept as valid the marginal significance of the nonsense result. Only the endogeneity-induced “bias inflation” perspective provides a reason why the nonsense result should not be accepted at face value: there is likely an omitted variable, such as “physical inactivity,” that biases the coefficients, and the high correlation between the body mass index and waist circumference inflates the bias as per the beta polarizing pattern.
Discussion, Extensions, Limitations, and Future Research
We demonstrated that VIF scores used as diagnostics for multicollinearity are at best meaningless and often misleading. This runs contrary to the two main viewpoints in academic research. The first, the threshold perspective, argues that low VIF scores provide a “silver bullet” to dismiss any and all multicollinearity-related concerns. Researchers near-universally interpret a VIF score below a threshold as providing support above and beyond the t-statistic for the validity of statistically significant results in the presence of multicollinearity. This perspective has become the default in current academic research, with the vast majority of multicollinearity discussions relying on its use. The fact that researchers are near-universally using VIF scores to buttress the legitimacy of statistically significant results, as per our methodological practice review, suggests that implicitly they are concerned with bias inflation and type 1 errors, rather than variance inflation and type 2 errors.
The second, less prevalent viewpoint, is the t-statistic perspective, which argues that VIF scores are superfluous to measures of statistical significance. The VIF scores are merely a component of the equation for variance, and therefore the t-statistic has already taken their size into account. This perspective relies on straightforward tests of significance to dismiss multicollinearity concerns.
However, multicollinearity remains problematic in organizational research, even in the face of below-threshold VIF scores and statistically significant coefficients. We have generalized the case of Kalnins (2018) to articulate the “bias inflation” perspective of VIF scores and multicollinearity. We showed analytically that omitting a variable inflates endogeneity biases and may cause polarized coefficients and type 1 errors. Using numerical examples based on the equations we have derived, we demonstrated that inflated coefficients often coincide with both type 1 errors and low VIF scores. Thus, both the threshold perspective and the t-statistic perspective are misleading approaches to diagnose and address multicollinearity concerns. Furthermore, we demonstrated that the endogeneity-based bias inflation may complement the variance-inflating effects of multicollinearity to increase the incidence of type 2 errors as well.
We provide two potential remedies for the multicollinearity problems identified above. First, we encourage the use of strong, exogenous instrumental variables. These will eliminate bias that results from omitting a variable and thus will render irrelevant the bias inflation effects of multicollinearity. Second, if instruments are unavailable, we encourage researchers to test only straightforward, unambiguous theory when using data that exhibits multicollinearity, and to ensure that correlated co-variates exhibit the expected signs and are of a reasonable magnitude.
We highlight some variants of OLS regression where the VIF scores might exhibit even more misleading behavior. First, the case of quadratic terms and interaction terms will often yield meaninglessly high VIF scores because of the often-high correlations between a variable and its quadratic or with an interaction that includes it, for example. The research methods community understands well (e.g., Edwards, 2001) that such correlations are benign in terms of bias and inference. The only bivariate correlations that will cause meaningful bias inflation and beta polarization among quadratic and interaction term coefficients are those between primary terms (Kalnins, 2024). Second, fixed effects models will also yield meaninglessly high VIF scores if the effects are included in the regression as dummy variables. These variants of OLS provide additional reasons not to use VIF scores, and thus to rely on simpler bivariate zero-order correlations to assess the possibly deleterious implications of multicollinearity.
We support the organizational research field's current practice of including full bivariate correlation tables so that readers develop informed opinions about the potential multicollinearity concerns. The correlation table should include all interaction terms and quadratic terms. Further, if a study relies on fixed-effects analysis, authors should present a correlation table for all variables with the fixed effects partialled out. The “within” bivariate correlations are of primary relevance and may be vastly different than those across all raw observations.
Our concerns extend beyond basic multiple linear regression to other circumstances that rely on similar mathematical assumptions, like multilevel modeling (e.g., Peterson et al., 2012). In multilevel modeling, OVB remains a concern for the accurate measurement of coefficients. Indeed, simulation studies have shown how multicollinearity in multilevel models can result in biased parameters and inaccurate SEs (Shieh & Fouladi, 2003). Because misspecification violates the same assumptions in these models as in multiple regression, multicollinearity among included variables is likely to produce the same sorts of errors as those presented here. Nonetheless, we will leave a definitive conclusion for future work.
A second area for future work is the incorporation of sensitivity analyses that attempt to determine whether an omitted variable is sufficiently strong to reduce or eliminate the statistical significance of a hypothesized result. For instance, the impact threshold of a confounding variable provides insight as to the correlation an omitted variable would need to exhibit with the dependent variable and with a focal independent variable to overturn support for a hypothesis (e.g., Busenbark et al., 2022; Frank, 2000). Similarly, Cinelli and Hazlett (2020) examine treatment effects and the necessary variance that they must explain for researchers to be assured of valid causal inference. And, turning attention to the degree of bias in the estimate, Oster (2019) formulates a proof to help calculate the precise point estimate given assumptions about the explained variance of an omitted variable. These sensitivity/robustness techniques may bring substantial potential benefits for researchers who face possible OVB and multicollinearity problems. However, Frank (2000), Oster (2019), and Cinelli and Hazlett (2020) are all based on extensive, subtle, carefully derived mathematical arguments. A thorough mathematical analysis would be necessary to clearly affirm or discount the use of these methods in our setting. We are interested in conducting this work in the future.
Finally, we note that the VIF score is strictly an artifact of the linear algebraic basis of OLS regression. Limited dependent variable models and survival duration models that rely on maximum likelihood for their solutions have no obvious equivalent of a VIF score. VIF scores simply cannot be reported for ML models. However, as reported by Kalnins (2018), simulations show that high zero-order bivariate correlations remain associated with bias inflation and beta polarization in these contexts. Type 1 errors driven by multicollinearity and endogeneity bias appear every bit as likely in these ML models as they are in simple multiple regression models.
We conclude by articulating the observation that, in organizational research, our counter-intuitive hypotheses are often subtle, and supportive results are not easy to identify as nonsense. We sometimes celebrate supportive results of counter-intuitive hypotheses for their novelty and use them as a basis to motivate new theorizing. Unfortunately, omitting variables leads to endogeneity biases and may lead to excess type 1 errors among included variables that are correlated with the omitted variable and with each other. Low VIF scores do not strengthen the legitimacy of counter-intuitive results with flipped signs or unrealistic coefficient magnitudes that appear to be statistically significant. Counter-intuitive results should be viewed with a skeptical eye because they are often in fact type 1 errors.
Footnotes
Acknowledgments
We would like to thank Myles Shaver and Evan Starr for helpful comments.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
