Abstract
In this article, we present and argue our assertion that current routine psychological testing of individuals is not valid. To support our assertion, we review the concept of ergodicity, Birkhoff’s theorem, and Molenaar’s manifesto, which together support our contention that the direct transposition of population estimations for producing inferences about the individual is not valid. We argue that this practice of direct transposition is the root cause of why routine psychological testing of individual is not valid. We then provide an example of a common application of psychological testing of an individual, explaining why this practice is not valid. Finally, we discuss how the intraindividual (or within-person) approach provides some prospect for valid individual testing and also introduces new challenges. We hope that our questioning of current psychological testing practices motivates researchers to propose and study novel methodological propositions to address the issues raised by our assertion.
Background/context
In this article, we present and argue our assertion that current routine or traditional psychological testing of the individual is not valid. We recognize that our challenge to conventional thinking and practice if accepted will have serious implications for psychological research and practice. Therefore, throughout this text, we seek to show the logical plausibility of our assertion, through a series of arguments. Before presenting them, we define what psychological testing of the individual is as well as what we call the routine or traditional practice of psychological testing of the individual.
Our notion of psychological testing of the individual encompasses any psychological practice that makes inferences about a particular person using one or more psychological tests. More precisely, we label or consider as routine or traditional practice of psychological testing of the individual, those specific situations when the psychological testing of individual produces inferences about the individual using the population estimations of psychological tests. We use the terms routine and traditional, because the use of psychological testing has become deeply rooted in the practice of psychologists since the early 20th century. To cite just one example of a longstanding practice that dates to the beginning of the 1900s, Intellectual Quotient (IQ) tests are commonly standardized to have a mean of 100 points and a standard deviation of 15 points based on a given normative population. Test scores are typically adjusted so that the score of a given individual is always compared to the mean, to the standard deviations, and to the percentiles of the normative population. In this traditional approach, usually the individual is tested just once. In conclusion, this practice is so traditional and at the same time so common in individual testing that it guides the underlying logic of construction and interpretation of test scores in Psychology.
Despite focusing on the routine or traditional practice of psychological testing of the individual, we want to be clear that we are not advocating limiting psychological assessment to the exclusive use of tests nor to the practice of comparing the individual to normative criteria. We understand that psychological evaluation is a broad process of inferences that uses information from psychological tests as well as other sources of information, such as the clinical interview, history information, observation, and qualitative data. In this sense, we emphasize that the scope of our article is limited to and focuses specifically on what we are calling the traditional or current routine practice of psychological testing of the individual and its implications for the evaluation and clinical practice in Psychology.
Psychological testing is widespread. It is used by companies and organizations, for example, in selection processes for employment and for promotion. In educational settings, psychological testing is used to assess students with behavioral issues or learning difficulties. It is used by government agencies in several countries for the licensure of drivers and those seeking to carry firearms. And psychological testing is used extensively as a forensic tool in the judicial and criminal justice systems. The advent of software applications which make it possible for subjects to more quickly self-administer psychological tests using tablet computers and smartphones portends an expansion of their use.
Finally, we note that the current routine practice of individual testing, in addition to being a century-old and rooted tradition, is an appealing and intuitive idea. Indeed, it is appealing precisely because it makes intuitive sense to generate inferences about a specific person using the population estimates. After all, if a person belongs to a given population, theoretically, estimates of the population could represent essential information about this individual. Our assertion rejects what is intuitive and thus goes against the “scientific common wisdom” that the traditional practice of individual testing is valid.
In subsequent sections, we present arguments supporting our assertion, which we acknowledge some readers will consider quite counterintuitive. We begin by presenting the concept of ergodicity and Birkhoff’s theorem, followed by Molenaar’s manifesto and discuss their implications for psychological testing of individuals. We conclude our article by introducing what we consider are several valid alternatives for psychological testing of the individual (with their own challenges), connecting them with the intraindividual (within-person) approach.
Concept of ergodicity
In the field of statistical mechanics, Boltzmann introduced the term “ergodic,” regarding his hypothesis: for large systems of interacting particles in equilibrium, the average time along a single trajectory equals the space average. For instance, in Physics and Thermodynamics, the ergodic hypothesis states that, over long periods of time, the time spent by a system in some region of the phase space of microstates with the same energy is proportional to the volume of this region, that is, all accessible microstates are equiprobable over a long period of time.
The ergodicity hypothesis as it was stated was false, but it generated intense investigation for the conditions under which it could be true. This in turn led to the birth of the concept of ergodicity as it is known today. Heuristically, we want to elucidate the conditions under which a process, say an ergodic process, will have, in a long run, an average along a path of one element evolving in time through the states of the process that is equal to the average of all possible states that the process may attain at any given time. This kind of process must be homogeneous along the time as well as throughout the space, and the necessary conditions for this to happen are that the process should not possess any absorbing state and any cycle and should be stationary. An absorbing state is a state where the process, once there, can no longer exist, and a cycle is a collection of states where the process circulates around without escaping. The definition of stationarity is presented later.
A modern description of the concept of ergodicity states that it is the study of the long-term average behavior of systems, with certain properties, evolving over time. The collection of all states of the system form a space X, and the evolution is represented by a transformation
A probability space is defined as (X, ∑, µ) where X is the sample space, or phase state, µ is a probability measure, and ∑ is a σ-algebra. A σ-algebra is a collection of all subsets of X that can be measured by µ, that is, subsets to which it is possible to assign a probability using µ.
Let (X, ∑, µ) and (Y, ϒ, ν) be probability spaces. A function
Let (X, ∑, µ) be a probability space and
This means that for the integers
It should be noted that
Now we may define ergodicity: Let T be a measure preserving transformation on a probability space (X, ∑, µ). The transformation T is said to be ergodic if for every measurable set E satisfying
Molenaar’s manifesto
In 2004, Molenaar published a manifesto, which stated that Psychology had made a historic mistake using population estimates to construct directly inferences about the individual. We note that the assertion being put forth in this article is an extension or a translational application of the central thesis of the Molenaar’s manifesto to applied psychology.
The ergodic theorems, as elaborated 85 years ago by Birkhoff (1931) and by von Neumann (1932) (through different approaches), provide the mathematical foundation for the concept of ergodicity. Molenaar applied an insightful strategy to develop his manifesto, which has been proved quite fruitful throughout the history of science: discern the relevance of the advanced knowledge of one field of science and then transfer it to another field. Molenaar recognized that the ergodic theorems in the field of statistical mechanics could be applied to psychological science since the concept of ergodicity should be applicable to any measurable process (Moore, 2015).
The ergodic theorems support the following. According to Birkhoff’s theorem, let (X, ∑, μ) be a probability space and
But what does this have to do with Psychology? The answer lies in translation of the theorem in terms of its implications in psychological testing. To do this, we need to return to the concept of ergodicity presented in the previous section. Let us then use the probability space defined previously where X is the space of all possible results of an IQ test, and the σ-algebra is the set of all sub-collections of those results. Suppose that N persons take this test, each one n times, that each test is repeated over a fixed and discrete period of time, and Ti is a rule that transforms one result to another one for the person i, and f is the total number of correct answers. Therefore, we have N processes varying in time in the same probability space. Could those processes be ergodic?
Using Molenaar’s notation, let
We acknowledge that our explanation up to this point has employed a considerable number of mathematical axioms. Let us then also present our argument from a semantic perspective. The ergodic theorems demonstrate that the occurrences of a system only exhibit the same behavior if they are ergodic. We may construe a system as a population, and the occurrences of the system as the members of the population. In this way, we can say that the ergodic theorems require that all members of the population must present a similar behavior or possess a similar structure if this population has ergodic characteristics. The behavior of individuals, being ergodic, is homogeneous and stationary: homogeneous because they present the same characteristic, and stationary as the homogeneous nature does not change over time. Assuming that a population possesses ergodic characteristics, if a test is valid for the population, then this test is valid for each individual who is a member of the population. Furthermore, the population estimates of this test should be directly applied to generate inferences for each individual of the population. Conversely, in nonergodic populations, the population estimates must not be transposed to generate direct inferences about the individuals of these populations.
Molenaar (2015) is blunt, arguing that ergodicity has implications for the entire scope of quantitative methodology in Psychology, which inevitably extends to encompass the construction of tests and their validation. The consequences of the classical ergodic theorems affect all psychological statistical methodology (Borsboom, 2005; Molenaar, Huizenga, & Nesselroade, 2003). Because a wide range of central psychological processes like learning, information processing, habituation, development and adaptation generally imply that some kind of growth or decline occurs, these processes are almost always non-stationary (violating the homogeneity in time criterion for ergodicity) and are, therefore, non-ergodic. This implies that their analysis has to be based on intra-individual variation in order to obtain valid information at the level of individual persons. (Molenaar, 2015, p. 37)
Evidence about the infringement of ergodicity in psychological constructs has existed for 20 years. Borkenau and Ostendord (1998) elaborated what may be the first empirical study of this violation. Collecting data over time from a set of people and applying time series factor analysis to investigate the factor structure of each individual, they found results that refuted the hypothesis that the five-factor structure or the Big Five Model of Personality should necessarily be found at the level of individual. Other empirical studies have followed. Among recent studies, two investigated the concept of ergodicity involving the Cattell-Horn-Carroll (CHC) model (Gomes, Araújo, Ferreira, & Golino, 2014; Gomes & Golino, 2015). In ways that parallel Borkenau and Osterdord’s conclusions about Personality, these refute the hypothesis that the CHC structure of abilities must necessarily be found at the level of the individual.
An example of the routine psychological testing of the individual
Although we have set out our arguments about the problem of the nonvalidity of the routine testing of the individual, we recognize that a concrete example of practical, day-to-day psychological testing of the individual has the potential to facilitate acceptance. Thus, we present an example, which illustrates the aspects that demonstrate the nonvalidity of routine psychological testing of individual.
Our example is the common situation in which a professional receives a patient for psychological evaluation due to learning difficulties in school. Through clinical interviews and observation, the professional gathers information to contextualize and understand the child’s problem. The professional must investigate the patient’s health history as well as the family, social, and academic context.
To investigate the reasons underlying the patient’s learning difficulties, the professional selects the WAIS-IV for the assessment of cognitive functioning. He chooses this instrument for the following reasons: (1) it is one of the instruments most frequently employed in clinical practice; (2) it is indicated for the age-group of the patient; (3) it allows the evaluation of several domains of intelligence, including general intelligence; (4) scientific studies report the association between intellectual performance and academic performance; (5) it has satisfactory psychometric properties at the population level which attest to the internal structure of intelligence dimensions as well as the expected association with external criteria (validity); in other words, internal consistency and temporal stability (reliability) have been established; and (6) the technical manual reports norms standardized on 2200 American individuals, ranging in age from 16 to 90 years.
Taking up the problem that underlies the evaluation process, the clinician intends to verify the level of performance of his patient in general intelligence, since one of his hypotheses is that the complaint of learning difficulties in school presented by his patient may be associated with a low level of intellectual ability. The clinician asks the patient to respond to WAIS-IV items only once. Based on the patient’s performance, the clinician constructs inferences about his patient by comparing his performance with the normative population estimate of the test. A key inference elaborated by the clinician involves the level of the patient’s general intelligence. The clinician infers that his patient has above-average general intelligence because the patient’s score corresponds to the 55th percentile of the normative population in WAIS-IV, indicating that the patient’s performance was equal to or greater than 55% of the individuals in his age cohort. For the professional, this inference about the patient’s intellectual functioning seems to be accurate and valid, since he was cooperative and motivated to perform the tasks of the test to the best of his ability.
The clinician takes for granted that the WAIS-IV is valid to measure general intelligence in the population, and thus, it is valid to use the instrument to make inferences about the general intelligence of his patient. The clinician assumes that he does not need to estimate the patient’s performance to make inferences about him; he can simply use the patient’s raw score obtained in a single application of the test and compare this score with the population estimate.
Continuing the example, the clinician also applies the MMPI-2 test once in this patient to assess his personality and interpersonal functioning. The clinician considers the hypothesis that the complaint of learning difficulties in school may also be associated with some personality trait of the student. He chooses to use the MMPI-2 because this test presents evidence of validity, reliability, and normative scores at the population level. The clinician uses the same approach to make inferences about the patient by directly comparing the patient’s raw score with the population estimate for the MMPI-2. Among the patient’s scores on the 10 basic scales, the raw score of 23 on the Depression Scale is considered important because it is high in relation to the normative population. This high score is associated with the presence of symptoms of depression, worry about the future, hopelessness, possible suicidal ideation, vulnerability, and feelings of isolation.
Note that the same assumptions underlying the application of WAIS-IV are repeated in the application of MMPI-2. The clinician assumes that if this test presents evidence of validity in the population, then it is certainly valid to evaluate the personality dimensions in his patient. Similarly, the clinician assumes that it is sufficient to use the patient’s raw score obtained in a single application of the test to make inferences about this individual by comparing the raw score with the normative population estimate of MMPI-2.
After the clinical interviews and the application of tests, the professional integrated the information to formulate a diagnostic conclusion and make recommendations. In a synthetic way, he concluded that the patient has the general reasoning and learning abilities necessary to fulfill the school’s expectations. Depressive symptoms seem to explain the learning difficulties in school. To complete the evaluation, the clinician developed a psychological report and provided feedback to the patient in a posttest encounter.
In this example, the nonvalidity of the routine testing of the individual according to the arguments based on the theorems of ergodicity is presented as follows: (1) the professional judged that the WAIS-IV and MMPI-2 tests are adequate to estimate the psychological characteristics of the patient. This judgment was based on the understanding that the existence of evidence of validity for these instruments, as reported in technical manuals and scientific literature, legitimizes the use of that estimate of the intelligence and personality of a particular individual. The clinician's only assurance that a test with evidence of construct validity at the population level equally evaluates the same construct in this patient is the common belief in the direct transposition of the population estimate to the level of the individual. However, according to ergodic theorems, this common belief is erroneous. Thus, it was a mistake to assume and directly transpose this evidence to the level of the individual, assuming that these tests and their items have validity to evaluate the intelligence and personality of that particular patient.
(2) The professional applied the instruments only once. The patient’s scores were compared with those of other individuals, according to the respective normative scale and standards of the tests. By using the patient’s raw score and turning it into a normalized score in the population, the clinician concluded that the patient has high average intelligence and high depression. In the case of intelligence, he identified that the true level of general intelligence should be between 106 and 120 or the 55th percentile. He followed the usual and traditional procedure of testing to perform a comparison of an individual’s scores with those of other individuals of the population to make inferences about a psychological characteristic of a particular individual. However, what assurance does the professional have that the patient’s intelligence and personality estimates are truly those obtained from a single application of each test? According to the concept of ergodicity, such inferences are not valid.
In summary, the problem in the example described earlier rested on the fact that the professional assumed that the tests are valid to report on a particular patient while they only reflect evidence about the population. In order to make an estimate at the level of the individual patient, it would be necessary to apply the test on successive occasions over time. Based on the scores obtained during these applications, it would be possible to capture the variability of the scores in order to investigate and confirm the construct validity. It is only after these steps that one could finally make the comparison with the population estimate. However, the practice of testing restricts the subject only to this inappropriate comparison with the population.
The psychological testing of individual implies the intraindividual approach
There are a significant number of studies that have considered the problem of ergodicity and have proposed methodologies to generate the estimation for the individual from the estimation of the population; Molenaar is coauthor of two. Here are five: (1) the general latent modeling approach (Adolf, Schuurman, Borkenau, Borsboom, & Dolan, 2014), (2) the integrated state-trait model (Hamaker, Nesselroade, & Molenaar, 2007), (3) the nonequivalence estimation method (Voelkle, Brose, Schimiedek, & Lindenberger, 2014), (4) the idiographic filter (Nesselroade, Gerstorf, Hardy, & Ram, 2007), and (5) the group iterative multiple model estimation (Molenaar & Nesselroade, 2014). Despite the integrative efforts mentioned, it is important to note that none of the above propositions was created to respond, theoretically or methodologically, to the specific challenges of the clinical testing of the individual. Thus, in methodological terms, a large gap remains to be bridged, which we will address and discuss more thoroughly ahead.
The clinical practice of individual testing is much more complicated than we would like it to acknowledge. To apply a specific test to a particular individual, the clinician will need to perform a validity analysis of the test, in order to verify if the test is valid for the individual. To do that, it is necessary to expand the testing of each individual to encompass an intraindividual approach.
This intraindividual or within-person approach is appropriate to make estimates at the individual level. Intraindividual analysis explores the individual’s own variance, that is, intraindividual or within-person variation. It is the opposite of the interindividual approach, which is focused on the variation among people or across individuals, which is necessary to be able to produce estimates at the population level, adequate to encompass and account for the range of individual differences.
Considering the routine psychological testing of individual, to date, we have been producing inferences about the individual using just the interindividual approach, which is not suitable for generating a valid estimation for an individual. In this context, if we want to use a correct approach, professionals need to change as fast as possible and adopt the intraindividual approach. As we stated previously, the interindividual approach is a powerful methodology and the correct strategy to estimate populations.
Why is the intraindividual approach the valid approach to estimate the individual? Why can’t we use the interindividual approach?
Any testing, whether at the population level or whether at the individual level, is based on the fundamental condition that the items of any psychological instrument capture some variance. Considering the interindividual perspective, it is the items of a psychological instrument that capture the variance, precisely through the difference of performance across individuals. Usually, the interindividual approach demands that the testing be applied just once on a reasonably large and heterogeneous sample of the population, increasing the chance that the variance can be captured and used for the correct estimation of the population. In contrast, the testing of the individual requires that the items of a psychological instrument capture the within-person. In this context, the variation across individuals is not at stake, but rather only the variation of the individual, on different occasions of testing over time. From the perspective of process, the intraindividual approach requires that the items of the psychological instrument be applied several times in the same individual, because the key is to capture the individual’s own variance. Only a test that is applied several times, which captures the individual’s variance, can permit the estimation of individual.
Let us now return to the methodological gaps, which we commented on briefly when we cited some methodological studies that sought to relate the individual estimation to the population estimation. To date, we have not encountered any empirical or simulated study which provides a reference or a guideline on what would be a reasonable minimal frequency of testing, enabling the individual’s estimate to be properly determined. Therefore, we do not know if we will need to ask an individual to perform a test 3 times, 6 times, 10 times, 30 times, or more, in order to be able to make valid estimates at the individual level. Furthermore, we do not know how many items a test would have to include in order to obtain a correct estimation for the individual. It is possible that the number of items applied causes some influence on the number of testing occasions. In addition, not just the number of items seems to be important. It is possible that the relationship between the ability of the individual and the difficulty of the items in the test might influence the number of testing occasions that is necessary to properly estimate at the level of the individual. If a test includes many items that are too easy for a specific individual, such items would not capture any variance, as it is likely that this individual would respond correctly from the first occasion of testing. Without any variance, these items would not contribute to the estimation of the individual, requiring either more items or more testing occasions. On the other hand, if a test includes many items that are too difficult for a specific individual to comprehend or answer correctly, perhaps this individual would fail these items, even after testing on multiple occasions. In this case, these items would not capture any variance too, similar to the scenario of the extremely easy items, leading to the same consequences.
So, if we already know about the necessity of repeating the testing to estimate at the level of the individual, then the optimum number of testing repetitions, the optimum number of items, and the optimum relationship between the initial level of the individual’s ability and the difficulty of items are unknown. In turn, this is a powerful obstacle. For example, if it is necessary that an individual performs the same test numerous times for their estimation, then in the clinical setting, the valid process of individual testing simply becomes impractical economically or logistically.
What then are the minimum conditions for estimation at the individual level?
If the clinician is evaluating a specific person, the key to obtaining the individual estimate is to know whether the test considered is valid for the population and is valid when applied to a particular individual. Considering this objective, we should consider that the clinician does not need to estimate precisely all the dimensions of that individual. It would be enough to obtain evidence that the test seems to possess a structure similar in the individual to that which possesses for the population. Note that possessing a “similar structure” is different than possessing the “same structure.” For example, at the population level, a test might measure only one dimension; while at a level of the individual, the same test could measure three dimensions. One of these three dimensions, however, should predominate and capture the greatest share of the individual’s variance, insofar as the identification of this specific dimension properly fits the performance of this individual. Therefore, in this example, the structure of the individual is similar to the structure of the population but it is not the same. As we noted earlier, the clinician does not need to estimate the entire structure of the individual, precisely and completely. Theoretically, it should be sufficient to find evidence that the structure found at the population level reasonably fits the individual’s data and the test seems to measure the same dimension at both levels (population and the specific individual).
Continuing the reflection on the optimal conditions for the individual testing, it is interesting to consider using the strategy to diminish, to the greatest degree possible, the complexity of the analyzed structure. We know that the simplest estimation of any structure is the analysis of a specific dimension, whether at the population level or at the individual level, because the greater the number of dimensions needed to estimate, the more data we will need to gather, and in the intraindividual approach it means more repetitions of testing on separate occasions. Therefore, more economical approach to the clinical context of individual testing involves using this specific dimension for the validity analysis (unidimensional model).
Under these conditions, if the clinician applies a test several times on a specific individual, and this test possesses evidence that it is unidimensional at the population level, it would be sufficient to know that the specific dimension fits the individual’s performance. It is important to say that the unidimensional analysis works well even in the case of tests that present multiple dimensions at the population level. In this situation, it is possible to select for the analysis just the set of items that are related to the target dimension, which the clinician is interested in evaluating.
To summarize our reasoning up to this point, reducing the complexity should be a key strategy to diminish the requirements for many repetitions of a test to obtain an estimation at the level of the individual. This approach is pragmatic and could dramatically reduce data collection requirements, implying fewer testing occasions.
How would this approach work?
Let us imagine that we intend to apply a general intelligence test to an individual. The test has 20 items and at the population level there is ample evidence of its validity and reliability for measuring general intelligence. So, what does the clinician have to do? In this example, he needs to obtain evidence that the specific dimension—at least a subset of items of this test—fits properly. It is not necessary that all the 20 items contribute to the specific dimension of the test at the level of the individual. As we noted previously, at the population level, the variation of the items is produced by different individuals. Conversely, at the individual level, the variation of the items is generated exclusively by the individual’s performance over the course of testing on different occasions. In this sense, we argue that some items that work well to measure general intelligence at the population level would not necessarily work as well for the estimation of a specific individual. Understanding this challenge, it would be reasonable and strategic to exclude some items of the test when estimating for a specific individual, as it would be understandable that the items excluded for estimating in the case of one individual might work well for another individual. Selecting items and searching for a salient dimension are just two interesting tactics that could be applied in the clinical context of individual testing.
Despite the substantial challenges related to dimensionality analysis in the estimation of the individual, there is a second question, which is complementary to the dimensionality analysis: the necessity to estimate the true performance of the individual. The traditional or conventional practice of testing usually asks that the individual perform a test just once because it assumes that a greater frequency of testing is not necessary to estimate the individual. We know, however, that it is necessary for an individual to perform a test several times to estimate the true performance of the individual. We have not encountered any reference or guideline regarding how to assess or measure the true performance of the individual. For example, we do not know if we must consider the rate of growth of learning that the individual exhibits when performing the test on several occasions over time. Instead of considering the rate of growth of learning, we could consider the performance achieved on the first occasion of testing; or the maximum performance achieved, or in the case that the individual whose performance possesses a logistic pattern where the individual improves his or her performance initially and then after some occasions there is no more relevant improvement, we would consider the asymptote. Ultimately, there are many methodological considerations that we would need to specify if we want viable and valid testing of individual. As we identified several intraindividual approaches, we have also called attention to a series of steps that would need to be performed properly. The challenges are substantial, as there is no established or traditional methodology to address the issues we have raised.
Conclusion
By affirming that the clinical practice of testing of the individual is predicated on the nonergodicity of psychological data, we do not mean to suggest that most professionals are cognizant of the ergodic principles as brought forth by Molenaar in his 2004 manifesto. In this spirit, we have the belief, indeed the conviction that when confronted with the concept of ergodicity and its implications, the clinical practice of psychological testing will have to undertake profound changes. Indeed this is one of our hopes regarding the impact and relevance of this article. We believe these observations about the implications of ergodicity require that professionals involved in the clinical practice of testing of individuals promptly address the issues raised.
As we have observed, the practice of inferring information about the individual from population estimates extends broadly, a consequence of the centuries-old practice in psychology of building theories, methodologies, and techniques of data analysis of population data that prioritize interindividual variations. In this sense, the now conventional practice of the individual testing is just one of many areas of psychology that was molded by the interindividual approach and the use of population estimates for generating inferences.
From the statistical point of view, in the interindividual approach “subjects are considered mere replications (e.g., interchangeable random extractions from the same probabilistic space having the same measure). This is expressed by the postulate that individuals are homogeneous in all relevant respects of analysis” (Molenaar, 2015, p. 36). Markedly, “the consequences of ergodicity affect the entire psychological statistical methodology” (Molenaar, 2015, p. 37), demanding a methodological revolution of intense breadth and depth.
The arguments about ergodicity, the properties of stationarity and homogeneity, and the problem of interindividual and intraindividual variation presented in this article derive from the Molenaar’s (2004) manifesto. This article recapitulates and explains Molenaar’s arguments and provocations and seeks to highlight the serious implications for the clinical practice of testing in psychology. We believe strongly that the flaws of current testing of the individual require that professionals move expeditiously toward more valid testing of the individual. The urgent need for change in the routine psychological testing of individuals employed in numerous setting is the singular message that we hope readers will take away.
In synthesis, this article presents mathematical formulae that sustain the formal conditions for the characteristics of all individuals of the populations to be equal to the characteristics of the population. Besides, those formulae were presented and discussed in the article as a formal approach to describe the ergodic concept and, mainly, to sustain our assumption which states that the traditional and current clinical testing of the individual is not valid.
Beyond the formulae presented in this article, which permit us to logically sustain our assumption that the current psychological testing of individual is not valid, as well as to sustain the conclusion that the psychological data, in general, are nonergodic, we cited a set of studies that corroborate this conclusion through simulations or empirical data (Borkenau & Ostendord, 1998; Gomes et al., 2014; Gomes & Golino, 2015; Molenaar, 2004, 2008, 2015). These studies estimated the individuals, and not the population, and found that the structure of the intelligence as well the structure of the personality which are present in all the estimated individuals were different from the structure of the population, refuting the traditional conception that the properties of each individual of the population have the same properties of the population. We make an allusion to the ideas of Karl Popper, which state that if only white rocks exist, then by empirically finding a black rock, you may refute the current theory, at least, ideally. The formal presentation of the ergodic concept that we present in this study, despite being sufficient, is empowered by the empirical findings that refute the current belief in the clinical testing context that the estimation of the population, and not the estimation of the individual, is sufficient to permit the construction of inferences about the individual.
In this article, we stressed the traditional clinical testing practice, especially regarding the inferences about the individual. The term “routine,” in the context of “Routine Psychological Testing of the Individual is Not Valid” does not mean an action taken regularly and over a period of time, which would erroneously permit the idea that the routine leads to the intraindividual measurement in the clinical testing. On the contrary, the term routine that we used in this article should be understood as a synonym of traditional and secular. In other words, we used the term “routine” to highlight the traditional, secular, and ingrained practice of testing the individual just once, assuming that the estimation of the population and the comparison of this estimation to the individual data permit to make inferences about this individual. This strong belief has historically produced an almost exclusive and common practice of cross-sectional data collection, which is deleterious to the estimation of the individual. As a consequence of this extensive and almost exclusive practice of cross-sectional data collection, traditional tests as WAIS, for example, do not collect empirical data which permit the estimation of each individual of the sample used in the validation studies of these tests. Unfortunately, as explained previously, to estimate the individual, it is necessary, at least for now, to collect many (30, 60, and 90) occasions of measurement for each person. Even traditional longitudinal data are insufficient to estimate each individual, since usually longitudinal studies collect a few observations over time, differently of time series, which collect many observations of each person over time (Molenaar, 1997). Considering this scenario, the inferences extracted from the validation studies of the traditional tests are all restricted to populations and not regarding the individual. For example, evidence of high longitudinal correlations or test–retest stability which comes from the estimation of the population only informs that the estimation on the population obtained similar results in two or more occasions, which are not the same information that permits to verify if each individual of the population has high correlations over time, which would indicate stationarity and homogeneity, and, for consequence, ergodicity. We must point it out that ergodicity concerns all the population, which, in our case, is composed by every individual with the same psychological properties and the process which is the test being taken repeatedly. It is impossible to prove ergodicity by looking at just one series, even a large series.
Finally, we make a point of the fact that we provide evidence, throughout the article, that, if we need to produce inference about a particular person, individual testing through time series is a superior approach in comparison to the traditional way of applying only once certain test in a particular individual. When we describe some examples of the clinical testing practice, our motive is only illustrative, regarding to show wrong common practices. In fact, we are not exhaustive at all in these examples and do not argue the detailed aspects of the practice. However, throughout the article, we sustain why the exposed examples are wrong practices to produce inference about the individual. We preferred this strategy: Simple examples, as illustrations, followed by an argumentation throughout the article, capable to sustain, logically and empirically why we state that these examples are wrong practices.
Highlights
The current routine psychological testing of the individual is not valid. The reason is the direct transposition of population estimates to the individual. Direct transposition is possible if the data present homogeneity/stationarity. Psychological phenomena are usually not homogeneous and stationary. An intraindividual approach is necessary to make estimates at the individual level.
