Abstract
A review of the methodological literature describing mixed-methods and quantitative and qualitative research paradigms suggest that though many have rejected the so-called paradigm wars there remains much focus on what is different about each research tradition. This has borne out in practice where professional organizations often have subgroups dedicated to the study of one tradition or another. Indeed, Human Resources Development Review has issued calls for manuscripts that explore this topic. This article examines the idea that there can be times when it is best to think of research as a monolithic paradigm rather than a distinct set of subparadigms. The reason for this is there are a number of common research scenarios where it is best to apply perspectives that might typically be characterized as qualitative as well as ones that are considered to be qualitative in orientation (e.g., using contextual information to make judgments about practical significance). There are other scenarios where the underlying goal of a procedure from, say the quantitative paradigm is similar to that from the qualitative realm (e.g., exploratory factor analyses and thematic analyses). There are of course real differences among the paradigms, but overemphasis on division might obfuscate how to conduct rigorous research. The article closes by encouraging readers to let their research questions dictate methodological approach, in the context of the purpose, rather than building questions around techniques that tend to align with different subparadigms.
Research endeavors in the social sciences contend with a common series of concerns such as the meaning and interpretation of findings, the degree to which such findings can be generalized to other settings, if results appear to be sensitive to alternative but reasonable analysis approaches, and so on. This point should strike any research methodologist as obvious, but what may not be so obvious is this observation applies to both the so-called qualitative and quantitative and mixed-method research paradigms. 1 This is noteworthy since these paradigms are generally treated as being distinct, and some have gone so far as to suggest that at least the qualitative and quantitative views are incompatible (e.g., Guba, 1990; Smith, 1983, 1984; Smith & Hodkinson, 2005). 2 Even though the incompatibility thesis and even the qualitative-quantitative paradigm wars have been rejected by many (Johnson & Onwuegbuzie, 2004; Newman & Benz, 1998; Ridenour & Newman, 2008; Tashakkori & Teddlie, 2003, 2010), methodological divisions remain in practice. Consider that in graduate school, methods coursework is often offered as a sequence of either qualitative or statistics tracks that sometimes meet under a mixed-methods banner, and there remains little literature on the explicit teaching of mixed methods (Christ, 2010; Creswell, Tashakkori, Jensen, & Shapley, 2003; Creswell, 2008; Johnson & Christensen, 2010). Organizations such as the American Education Research Association and American Evaluation Association continue to include specialized groups that focus only on one paradigm, and funding streams will at times hint at a desire to use only one approach. Indeed, paradigm differences of this sort can still go so far as to contribute to flux in professional identity and controversy, as has recently been seen in the field of anthropology (Wood, 2010). This overall concern is consistent with a call made by the journal Human Resources Development Review that asks for further discussion of quantitative, qualitative, and mixed-methods development (Reio, 2010).
There is furthermore a wide literature base that describes differences between major paradigms of postpositivism, constructivism, and pragmatism, with associated discussion of methods that belongs within each camp, which are, respectively, thought of as quantitative, qualitative, and mixed methods (see Denzin & Lincoln, 2005, for a description of how qualitative and quantitative research differs; see also Becker, 1996). They argue:
quantitative research relies on and accepts positivism and postpositivism whereas qualitative research accepts postmodern sensibilities (the authors acknowledge this is not true of all qualitative researchers);
qualitative researchers tend to think they can better capture the individual’s point of view via interview and observation, whereas many quantitative researchers view such techniques as unreliable and subjective (or perhaps Denzin and Lincoln are arguing that the way in which these techniques are used is viewed as promoting subjectivity);
qualitative research examines the constraints and realities of everyday life whereas etic science is based on probabilities and randomization, and these features fall outside the constraint of daily life;
qualitative research focuses on securing rich descriptions of data and quantitative researchers are less concerned with detail (see Denzin & Lincoln, 2005, pp. 10-12).
Although there is utility for making methodological distinctions such as these, such descriptions may also promote unnecessary division, and this can in turn undermine understanding of social science research design. Consider that experimental design literature extols the virtues of interviewing to try to understand what research participants in different study conditions may be thinking (e.g., Shadish, Cook, & Campbell, 2002) and qualitative research, such as those who describe ethnographic design, will describe how statistical data can be useful for understanding cultural issues (e.g., Schensul, Schensul, & LeCompte, 1999). As another example, when it comes to case studies, which are typically classified as qualitative research (e.g., Gall, Gall, & Borg, 2006), Stake (2005) states that these designs may have qualitative, quantitative, or mixed-method orientations depending on the researcher’s questions. In terms of focus on a single person, look no further than the single-case/single-subject research literature for examples of how idiographic causal inference, which focuses more on the individual (as opposed to average effects). Furthermore, single-case designs are typically conducted within real-life circumstances and observations are often the key method for gathering data (Kratochwill & Levin, 1992). In sum, notions such as quantitative researchers eschew interviews, qualitative researchers do not utilize statistical information, and so on, are too general. Given the typical distinctions between quantitative and qualitative paradigms are generalities, we argue it is better to focus more on the point that the purpose of the research is what drives the method and there must be a logical connection between the two. If, for example, one has some idea that is testable and is interested in generalizing findings, then it would seem there is a quantitative, postpositive orientation at hand. It may well be that numbers and the use of probability theory can be used to pursue the question, but certainly interviews and observation can and often do serve as useful supplementary methods. Indeed, some have argued qualitative techniques can even, by themselves, generate causal data (Maxwell, 2004; Shadish et al., 2002) even if they do not rely on frequentist probability statements. By contrast, if a research question focuses on whatever meaning study participants construct about some event, then a more qualitative and constructivist perspective can help. Yet there is no inherent reason why qualitative themes cannot be used to generate surveys in order to more efficiently gather data from larger samples, and this can open the door to a number of statistical techniques.
The purpose of this article is to review a series of common research issues to show there is nothing inherently quantitative, qualitative, or mixed about them but that, instead, they are better understood by thinking first that research itself as a monolithic paradigm that can at times be splintered into subparadigms. This is not to suggest research methods should not be classified into varied typologies with different goals. Such tasks can help express the goal and scope of work, and thereby, facilitate understanding (Nastasi, Hitchcock, & Brown, 2010), but there can be value in first understanding what paradigms have in common. At the outset, it is important to recognize the philosophical debate that parallels the paradigm wars and whether such wars were even worth having or continuing. Other sources cover this discussion well (e.g., Ridenour & Newman, 2008; Tashakkori & Teddlie, 2003, 2010), but we wish to circumvent it by acknowledging that the authors admittedly hold a largely postpositive view (the taller of the two authors might espouse a more pragmatist perspective but has yet to figure it out). Competing philosophical frames can provide a home for some questions that, for whatever reason, might obviate desire to generalize findings or argue for their transferability to other settings. Our article does not address such studies and refer readers to Denzin and Lincoln (2005) for a contemporary discussion of their application. Another caveat is the breadth of research questions prevent complete discussion of all design concerns that might come to mind. In lieu of providing an exhaustive review, we only wish to give voice to the idea that qualitative, quantitative, and mixed-methods research are often not as different as many appear to think and, whatever the prevailing paradigm, solid design requires close linkage between question and methods. 3 So only a few topics are discussed below. These are
interpreting statistical significance and replicability of findings,
practical significance and the need to understand context,
how data saturation is connected to generalization and sample-size concerns,
triangulation and negative case analyses,
the similarity between reliability and external auditors,
thematic analyses and factor development,
use of qualitative techniques to better understand participant section in quasi-experiments.
Interpreting Statistical Significance and Replicability of Findings
Classic Fisher/Neyman–Pearson Null Hypothesis Statistical significance testing (NHST) arguably lacks any qualitative decision making because of its reliance on comparison distributions and mechanical procedures. Interestingly, this procedure has also been heavily criticized in the literature (e.g., Schmidt & Hunter, 1997) and there was talk of even banning NHSTs (see Kline, 2004; Thompson, 1998, 2002a-b; also Wilkinson and Task Force for Statistical Inference, 1999, for overviews of related history). The key complaints about NHST are the approach is often misunderstood (Cohen, 1994) or thought by some to promote rote, even thoughtless, conduct of research (see Morgan, 2003). Despite these concerns, cursory perusal of contemporary empirical literature will show NHST has not been banned, and indeed, is still regularly used. This is largely because anytime one wishes to infer observations from a sample are reflective of some underlying population, then there should be some accounting of sampling error. Failure to do so leaves reasonable concerns that observed sample characteristics are not reflective of populations from which they were drawn. NHST provides a method for ruling out sampling error that, in the context of the approach, is viewed as being due to random, or chance, variation.
Key decisions in the process is understanding the probability of rejecting a hypothesis of no difference, when there is no true effect in an underlying population, relative to failing to reject the null hypothesis when one should have. In other words, balancing the probability of Type I and Type II errors entails a key judgment. We contend that the selection of the alpha level (.10, .05, .01, etc.) is generally more of a qualitative judgment than a statistical one. One should ask themselves why they prefer a given alpha level? What are the costs of making a Type I error versus the cost of making a Type II error for the stakeholders in their particular setting? For example, are higher Type I cutoffs acceptable in the case of efficacy studies that are designed to establish impacts under lab like settings as opposed to real-world effectiveness trials where results should be able to inform wider program adoption? We think it is likely that alpha levels are chosen without giving this much thought and the selection of a .05 p value is just convention (Kline, 2004). Decisions to be made about management training programs, product adoption, pharmaceutical choices, and so on, offer direct examples of this issue. The efficacy of an innovative training program might, for example, necessitate less stringent evidence given the relative risk of a Type II error, especially since supportive evidence would not suggest the program be adopted but rather further studies. By contrast, if the goal is to establish a management training program is effective even under real-world circumstances, then a decrease in Type I error would on balance seem to be the desired focus.
The smaller the calculated p value for a given cutoff alpha, the more likely the results will be replicable (at a specified alpha). This merits some discussion since it is possible that a study could be statistically significant and have a “large effect size” but not be replicable (see Figure 1). Figure 1 is meant to be a common depiction of a result from a design with a single treatment and control group. In this case, the treatment impacted the mean of the treated group so that it is 1.96 standard deviations above the mean of the control group. In this case, a null hypothesis is rejected using the standard critical value for a two-tailed test (i.e., the results are statistically significant at a .05 level). Assume the sample of the treated group well represents the treated population. Now consider what can be expected if additional samples from the treated population are obtained and compared to a control group. Holding all things equal, 50% of the treated samples will have a mean that is less than 1.96 standard deviations above the control; that is, if the treatment sample is representative of the true treatment population, and we sample again from this population, then 50% of the time results will be replicated. Therefore, if one is going to sample from the treated population, 50% of the samples would not have means that are large enough to reject a null hypothesis (assuming n > 12). Therefore, only 50% of the time will statistical significance be replicated. If the observed p value was .01, the treated mean is about 3 standard deviations above the mean of the control, and the critical cutoff is the same (.05), subsequent samples from the treated group would yield means that would reject a null hypothesis 72% of the time. If p =.001 then 90% of subsequent random samples from the treated population would yield means large enough to reject the null hypothesis (Newman, McNeil, & Fraas, 2004). More conventional thinking about study replicability is seen in the meta-analysis literature where an overall aggregated effect, weighted by the sample sizes in each study informing the statistic, is developed. Confidence intervals around this aggregated figure can be generated and it should be the case that the interval will be smaller than component studies as a function of the aggregated sample size (Grissom & Kim, 2005). This feature well establishes the point that more information is simply known as a function of replication.

A demonstration of the concept of replicability
It is the very notion of replication that invokes a need to understand the context of each study. If one should be cautious when making inferences about a single study, then it becomes important to understand how treatments in prior studies were implemented, whether these applications transfer well to a new context, and overall whether there is simply enough evidence to warrant adoption of some new practice. This is especially true when making judgments about causality as related reasoning is explicitly described as qualitative in nature (Shadish et al., 2002).
Practical Significance and the Need to Understand Context
We believe NHSTs play an important role, but they also play a small one when making decisions about the stability and importance of a finding. The limitations of NHSTs are well documented in the literature, but much of the focus has been on supplementing them with additional statistical procedures and paying careful attention to prior work. This is, of course, sensible advice, but we also think that qualitative research can and should play a vital role when making judgments about the practicality of findings, and whether some new practice should be adopted. Indeed, it seems like an unnatural division to espouse only one research paradigm when making decisions about program adoption. It seems more realistic to tackle the various aspects of program efficacy, effectiveness, adoption, and generalization, and social scientists are better served by thinking along a continuum of qualitative and quantitative inquiry.
Given NHST limitations, the prevailing point of view held by most methodologists is they should be supplemented, if not supplanted, via construction of confidence intervals (which necessarily encompass decisions to reject a null hypothesis, or not, at given alpha levels) and effect sizes (Callahan & Reio, 2006; Kline, 2005; Thompson, 2002b). These additional steps should minimize another source of confusion about NHST; that is, results from this class of analyses are conflated with sample size. To clarify, very large samples can render statistically significant results even when whatever was observed in the sample is entirely unimpressive. Likewise, potentially interesting findings may not be found to be statistically significant in the context of small samples. Recall Type II error is the scenario where a null hypothesis was not rejected but should have been; ceteris paribus studies with small samples run greater risks of this type of error. This difficulty in interpreting NHST has given rise to the notion of practical significance (Kline, 2004). Huck (2009) states practical significance is “focused on a study’s possible impact on the work of practitioners and other researcher . . . the determination of practical significance necessarily involves anticipating people’s subjective reactions to study findings” (p. 228). Practical significance is of interest here because of the subjective element of the idea. Researchers can and should consider research context, practitioner perceptions about issues like the size of impacts, feasibility of implementing some new treatment, and the social validity (Messick, 1995) of measures used in the study. To understand these issues well, researchers might interview key informants, invoke the notion of prolonged engagement in a context, and so on, to judge practicality, and these techniques are of course aligned with qualitative research.
Another approach where knowledge of what practitioner’s think would be practically significant could be the use of non-nill nulls (Cohen, 1994). Obviously a true null hypothesis (a nill null) states there is no relationship—a zero relationship—between whatever variables are being studied. In the simple case where one is testing some intervention impact with a treatment and control group, the mean of one group minus the mean of the other should equal zero under the null hypothesis. This is reasonable in a setting where random assignment is used because the allocation process should equate groups (on expectation) so posttests should be equal in the absence of a treatment effect (Shadish et al., 2002). A non-nill null, on the other hand, can test nonzero relationships between two groups, and this can make sense both in quasi-experimental settings where equating should not necessarily be expected and in cases where a researcher might wish to apply specialized hypotheses to accommodate what a practitioner might find to be useful. For example, suppose a researcher is examining the treatment impact of some novel dieting scheme. A nonzero difference between a treated and control group may not make much practical sense. But if a dietician would be interested in a treatment effect of, say, a 5-pound difference, then a non-nill null that accommodates such a test could be helpful. Such an analysis could in essence be done by subtracting 5 pounds from everyone in the treatment group and testing the resulting difference between the two groups in the traditional manner. Rejecting a null hypothesis would then entail saying there is at least a 5-pound difference between two groups, instead of the more typical application that examine whether there is a zero difference between the two groups (to clarify, this would be a null hypothesis test, but all scores are adjusted for a constant). An alternative, and far more common idea, is to set a minimum detectable effect size to accommodate a 5-pound difference between groups and power a study accordingly. Here, retaining a null hypothesis might mean the diet had an effect, but not one that is powerful enough to be deemed practical.
In either case, one needs to identify what a practical difference would be. Why 5, and not 7 or 10 pounds? One has to a priori identify the minimum cut score difference that has practical significance. Fraas and Newman (2000) worked in a cardiac rehabilitation setting went to Weight Watchers to determine the minimum weight loss that was found to show a pragmatic effect. The 5 pounds had a physiological difference on heart rate, blood pressure, and blood sugar. The problem they found is that it is often difficult to a priori identify a meaningful difference. There may need to be prior research that is done to determine the cut score that would indicate a meaningful difference, and that difference may vary by reading levels and for different students in a variety of settings (interaction vs. main effects). Identifying the relevant cut scores may be valuable, but they are not easy to obtain. But doing interviews with key stakeholders may help uncover information needed to judge practicality.
On the other hand, we argue that one should establish statistical significance (i.e., observed differences are not due to chance) before inferring practical significance. This is because if it is reasonable to assume an observed difference is due to chance, it becomes difficult to assume it is practical. Indeed researchers are left with the old problem of possibly making a Type II error in this case, where one is tempted to argue a difference is meaningful but cannot rule out sampling error. The key point here is interpretation of NHSTs takes on qualitative undertones. Thompson (1998, 2002a-b) and Rosenthal, Rosnow, and Rubin (2000) would, reasonably, ask that the magnitude of an effect be judged by extant literature on the given topic and nature of the outcome measure (which strikes us as a value-laden judgment). We suggest researchers routinely go beyond the literature when judging practicality and should ponder the nature of the outcome, cost of treatment, whether it is feasible to implement, and so on. Now we turn to a new topic.
How Data Saturation Is Connected to Generalization and Sample-Size Concerns
Qualitative research, like quantitative research, must consider how a sample has been collected, the size of the sample, and the consistency of responses elicited from the sample all while not losing sight of the questions at hand. The concept of saturation, though used predominantly in qualitative research, is in essence based on an underlying assumption that sample data are reflective of a population of responses (e.g., interview responses). The idea here is that when the level of saturation is truly reached, no matter how many additional interviews are conducted, no new concepts should emerge. From a postpositive point of view that relies on inference, if saturation really occurs, one can generalize from the relatively few interviewees, generally 6 to 18 (Guest, Bunce, & Johnson, 2006), to the entire populations that the respondents they are supposed to represent. Such a sample size is generally much smaller number than most quantitative researchers would be comfortable with, and there may be sampling error that needs to be considered, but these concerns are generally not acknowledged in most qualitative articles as far as we can tell.
Qualitative researchers should, however, be generally concerned about the degree to which a sample is reflective of some phenomenon of interest, and this has implications both for sampling procedures and sample size (Onwuegbuzie & Leech, 2007). Even in the case where only a single participant is of interest, one can ponder whether some set of responses (say to interview questions) are reflective of what a respondent might continue to say if given a chance to speak more. If, for example, a researcher asks a manufacturing plant manager about his or her feelings regarding the use of metal detectors to promote plant security, most researchers would take umbrage if the interview happened just after a violent event occurred (unless the research question fixes chronological concerns to a set point in time, such as examining how a manager feels about metal detectors after experiencing trauma). There are additional assumptions here; for example, respondents could have offered radically different insights if asked different questions, may have responded differently if a different interviewer elicited responses, and so on. Arguing that saturation has been reached is an assumption that these are not plausible concerns. Of course, there is also the more obvious case that responses from a small group of people reasonably generalize to some larger group. Here there is simply an assumption that interviewing new people will not seriously alter conclusions based on the dataset.
Onwuegbuzie and Leech (2007) use the term qualitative power analysis to call attention to these issues. They discuss the need for qualitative researchers to discuss how samples were collected. Random sampling can be applied in qualitative settings, though the central concern of determining whether a sample is reflective of some underlying population certainly is within the realm of a qualitative methodologist. In cases of nonrandom sampling, Onwuegbuzie and Leech note there is discussion about how to maximize the capacity to generalize observations to some larger population of possible responses. In the cases where sampling techniques undermine generalization, these authors recommend careful reframing of research questions, study parameters, and discussion of limitations to promote defensible interpretation of data. These concerns are old hat in a typical discussion of “quantitative” designs.
There are, of course, designs where it may be possible that the entire population of interest is available to the researcher, eschewing the need for sampling participants (though the above concerns about sampling a set of responses from all possible participants remain). This obviates any particular need to apply classic, inferential, statistical analyses. Although such studies are generally thought of as being qualitative, they are not necessarily so. As noted earlier, single-case (i.e., single-subject) designs do not generally endeavor to control for sampling error, and, again, some case studies can make use of quantitative data.
Triangulation and Negative Case Analyses
Qualitative researchers define triangulation as a search for converging evidence from multiple data sources, data points, investigators, theories, and methods (Brantlinger, Jiminez, Klingner, Pugach, & Richardson, 2005; Nastasi & Schensul, 2005; Patton, 2002). The capacity to explain triangulation when it occurs and when it does not occur (it may not surprise researchers when different stakeholders disagree on a topic) can promote validity, but Maxwell (2004) reminds us that this methodological approach does not guarantee triangulation will do so. After all, disparate data sources such as interviews, records, and focus groups all deal with the pitfalls of self-report and different stakeholders may be biased in the same manner about some topic of interest to the researcher. Nevertheless, the technique is widely written about and appears to be a mainstay of most qualitative inquiry interested in examining validity of findings. Getting back to the primary point of the current article, this is another way of saying that data may be reliable but conclusions drawn from them are not necessarily valid.
The notion of triangulation is also not solely in the province of qualitative methods (Bazeley, 2003; Campbell & Fiske, 1959) and cross-method triangulation is a key concept in mixed-methods work (Creswell & Plano Clark, 2011; Johnson & Onwuegbuzie, 2004). In the statistical realm, particularly in sensitivity analyses, whether different but reasonable analytic choices yield comparable findings is a key technique. Furthermore researchers will routinely consider whether findings hold across different measures, which is akin to triangulation across different data sources described earlier, and the very notion of triangulating findings with theory shares much in common with confirmatory analyses. In sum, triangulation is global in nature, as it encompasses many ideas and is an excellent demonstration of the point that different methodological traditions hold much in common.
The idea of negative case analyses/sampling is where qualitative researchers attempt to locate and account for discrepant and/or disconfirming evidence (Brantlinger et al., 2005; Nastasi & Schensul, 2005; Patton, 2002), that is, finding that evidence runs contrary to overall findings and then doing the work needed to explain such evidence or at least recognize its presence as a study limitation. As with the case of triangulation, this is not an exclusively qualitative concept. Consider that, in a quantitative sense, the negative case may be considered to be an outlier. Traditionally, an outlier is an infrequently occurring score that is approximately 2 or more standard deviations removed from the average score for that group. The first thing that needs to be determined when considering a deviant score is whether it is a measurement error or an actual occurrence. If it is judged to be a measurement error it is disregarded. If it is considered not to be a measurement error, one has to make a judgment based on its frequency of occurrence, the potential insight one can get from considering the deviant score in the analyses, and the purpose of the research. If the purpose is to be able to predict an outcome better than chance, and the researchers would be right 99 times out of 100, it may be cost-effective to disregard the outlier. On the other hand, if the purpose is to establish a causal relationship and to understand the underlying dynamics, it may be necessary to consider the deviant score in the analyses; that is, in doing qualitative or quantitative data analyses, one should not mechanically decide to implement rules about negative cases or outliers without considering the broader implications. In the end, the thoughtful and transparent consideration of all available data, regardless of paradigm, is a hallmark of good research.
The Similarity Between Reliability and External Auditors
The broader concept of reliability is a central concern to quantitative research. Of course in a psychometric setting the notion of score consistency is a necessary element to establishing validity (Crocker & Algina, 1986). In observational settings when researchers are making inferences about the presence of operationally defined behaviors it is critical to establish interobserver agreement (Horner, 2005; Kratochwill & Levin, 1992). Following the theme of this article, it is also the case that qualitative researchers often need to determine whether there is interobserver agreement when it comes to establishing their own findings. Brantlinger and colleagues (2005) describe the use of external auditors to examine if inferences and themes are logical and grounded in findings. Nastasi and Schensul (2005) point to a similar concept when referring to referential adequacy. These ideas overlap with the above discussion of validity, as is the case whenever reliability is discussed. The specific focus on reliability we mention here pertains to the narrower idea of perhaps seeing whether external auditors would come up with similar themes when looking at a portion of raw data or similar conclusions. Like interobserver agreement, one can wonder whether findings are consistent across researchers.
Thematic Analyses and Factor Development
Yet another area in which there are conceptual overlaps between qualitative and quantitative research is thematic development. Thematic development in qualitative research is similar, in many respects, to the identification of factors in exploratory factor analysis. Typical qualitative analyses entail identification of themes (conceptual factors) that emerge from the data and determining whether these themes are consistent within and across individuals (Patton, 2002). In the context of factor analyses, a factor consists of two or more items that appear to be measuring an underlying concept or construct (Darlington, 2002; Kline, 1994; Thompson, 2004). A theme may be considered to be similar to a factor in that it is conceptualized by identifying a number of statements that appear to cluster together to form a meaningful underlying concept or construct (Hitchcock et al., 2005). In exploratory factor analyses the factor loadings are used to in essence make a subjective judgment about the relative importance of that item to the factor, thereby aiding in the naming (interpreting) that particular concept or construct. There are a number of guidelines to judge the number and stability of factors (Thompson, 2004), but, in the end, it requires researcher knowledge of the structure of measures and overall knowledge of the topic to name and interpret a factor. A good example of mixing qualitative and quantitative techniques, in a multivariate sense, can be seen when using Q methodology. Q methodology consists of basically two procedures: a focus group that subjectively identifies the concourse of items, and a Q-factor analysis that quantitatively identifies the profiles or the types of people that emerge from their evaluation of the item concourse (Newman & Ramlo, 2010).
Frequently, when a confirmatory factor analysis is computed, it is done to determine how well the empirically emerged factors fit into an a priori theoretically identified framework. This is a deductive analysis. However, if one did a factor analysis to determine the underlying factors by seeing what emerges from the data, without any predetermined theoretical frame, this would be a more inductive process and would be more consistent with the underpinnings of a qualitative analysis. In addition, exploratory factor analysis tends to be sample specific (Thompson, 2004) and can be considered in a qualitative sense to be descriptive and heuristic. Again, one has to see that it is not the quantitative label or analytic procedures that makes something qualitative or quantitative, but it is the underlying assumptions that one makes about the data. To force things into a dichotomy (i.e., qualitative or quantitative), instead of a continuum, without considering purpose, intent, and assumptions, can be misleading.
Use of Qualitative Techniques to Better Understand Participant Selection in Quasi-experiments
Whenever random assignment is not feasible researchers can use some other selection method for establishing treatment groups to examine causal effects for a given intervention. Nonrandomized trials, or quasi-experiments, are thought to yield inferior causal validity (with perhaps the exception of a regression discontinuity design) because some unknown variable may be responsible for the selection into treatment conditions (Shadish et al., 2002); that is, whereas randomized trials can be expected to equate groups on all variables, whether observed or not, quasi-experiments cannot make the same claim. Of course, quasi-experiments range from ones that can hardly make any empirical statement that groups are equivalent at baseline to sophisticated designs that can argue for group equivalence on dozens of observed variables, or even more, using techniques such as propensity score matching (e.g., Hong & Raudenbush, 2005). In the case of almost all quasi-experiments (some may say all depending on how convinced they are that a given matching effort yields convincing evidence of equivalence), the basic issue is the selection model, that is, how units were selected into treatment conditions, is at least partially observed and not perfectly understood (Shadish et al., 2002). For this reason, it can be advisable to use qualitative methods to help understand the selection process. Assume one has a retrospective trial where units have already been assigned to conditions. Interviewing study participants and other stakeholders about the selection process may yield powerful insights into study group formation.
Our purpose here is to again argue that the walk is often not so far between qualitative and quantitative methods. Maxwell (2004) points out that qualitative inquiry does ask causal questions from time to time. If the basic premises for qualitative argument are applied (i.e., the causal variable must precede the effect, the variables in question should co-occur and there are no plausible explanations that undermine a causal argument) we see that there is no inherent need to apply inferential statistics though these certainly help. It is not difficult to imagine qualitative case studies and other designs that do ask a causal question (again see Maxwell for examples) and then it becomes easier to see how research does take on some monolithic characteristics.
Conclusion
These issues discussed here touch on the intimacy between the researcher and the participant, how researchers construct reality as they interact with data, make value judgments and interpretations, measure phenomena using techniques that are typically thought of as qualitative (thematic development) and quantitative (factor structures), and simultaneously deal with variables and social processes. Our primary hope is to impress upon readers the notion that we may be better served by thinking first that research is research and any paradigm selection beyond that is a more ancillary concern. Divisions between quantitative, qualitative, and mixed-methods approaches are arguably reified more by a need to label approaches than by true differences in purpose. Overemphasizing such differences can yield false notions such as qualitative studies cannot yield causal conclusions, quantitative work eschews exploratory analyses, triangulation is a qualitative technique, qualitative research does not use statistics, and so on. More critically, agreeing with the methodological division narrative can yield confusion since research purpose should dictate technique, and failure to see this can perhaps undermine research quality (Ridenour & Newman, 2008). If one believes that a single paradigm can lay claim to research purpose, this may contribute to a false sense of ownership of strategies and goals, which in turn may promote impoverished design and blind researchers to directions in which they may take their work. In short, the authors of this work suspect that labeling left over from the so-called paradigm wars, or perhaps just due to a need to organize, may be contributing to confusion about what sorts of research questions can be asked in one paradigm or another, as well as contribute to confusion around what techniques might be used. This can inhibit the development of teaching and the training of people in the social sciences to be researchers who can let their questions dictate their techniques and methodologies, rather than have their techniques limit their questions.
Footnotes
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
The author(s) received no financial support for the research, authorship, and/or publication of this article.
