Abstract
Psychology has at least three explanatory problems: (a) it continues to form and promote separate schools and camps that mainly work in isolation from each other or against one another; (b) each camp advocates its own vocabulary and set of basic concepts thereby further fractionating the professional identity of psychologists, explains small domains of interest to group members, but cannot provide an overarching perspective capable of transitioning psychology into a mature science; and (c) each minitheory promises mechanism information but only provides functional explanations that impute causality. A potential solution to our explanatory problems is presently available but has not yet been published in an organized way because investigators have not fully appreciated the broader implications of their work. This article aims to articulate those broader implications for psychological science in the form of a network learning theory perspective that provides a common vocabulary and set of three core and eight corollary empirically supported psychological principles. The adequacy of this approach has been proven mathematically.
Psychology has at least the following three major explanatory problems that have kept it from developing beyond its current preparadigmatic state (Kuhn, 1996).
First Explanatory Problem
The first explanatory problem is that psychologists have formed separate schools and camps each with its own vocabulary, theoretical orientation, methods, findings, and adherents since its formal inception in 1895, resulting in what Kuhn (1996) characterized as preparadigmatic science. Mischel (2008) described this theoretical diversity as a consequence of our “toothbrush problem”; “Psychologists treat other peoples’ theories like toothbrushes — no self-respecting person wants to use anyone else's” (p. 3). Kruglanski (2001) referred to our “theory shyness” (p. 871); our reluctance to move beyond midlevel theorization to more comprehensive explanations. Gigerenzer (2009) observed that “much of the theoretical landscape in psychology resembles a patchwork of small territories” (p. 22) that substitute surrogates for theory in the form of one-word explanations, circular restatements, and lists of dichotomies. Such theoretical diversity has been interpreted as a sign of scientific health by some (McNally, 1992) and a crisis with corrosive properties (Staats, 1983) that threatens to tear psychology apart (Spence, 1987) by others. Science values parsimony. All else being equal, we prefer simpler interpretations. Parsimony also extends to theory. All else being equal, we prefer to use fewer theories to explain our facts. Ideally, one theory is best. Introductory psychology textbooks demonstrate that we are far from this ideal. The first goal of this article is to introduce an intellectual platform with a common vocabulary and set of core and corollary principles that can transition psychology into a mature and consilient (Wilson, 1998) science. We currently need but do not have a uniform way to understand psychological phenomena.
Second Explanatory Problem
The second explanatory problem is that the theories and models that we do have are narrow and limited to a small part of psychology. For example, the Rescorla-Wagner model of classical conditioning (Rescorla & Wagner, 1972) has been successful within the limited domain of classical conditioning; however, classical conditioning is a very minor part of psychological science. The Rescorla-Wagner model of classical conditioning does not even generalize to operant conditioning, which was once thought to be an entirely different phenomenon, let alone generalize to other areas of psychology. Festinger's (1957, 1964) theory of cognitive dissonance was highly successful within the subfield of social psychology but it was limited to attitude formation and change. Even during its heyday, Festinger's theory was not used in other areas of social psychology, let alone all other areas of psychological science. It no longer drives nearly so much research as it once did and thus, its present impact is more historical than contemporary. Other examples of minitheories that have a limited scope can be found in chapters of introductory psychology textbooks. We currently need but do not have any overarching theory that pertains generally to psychological science.
Third Explanatory Problem
The third explanatory problem is that psychology lacks mechanism information, including a lack of consensus regarding what constitutes an acceptable mechanism. Psychologists frequently use the term “mechanism” but ultimately end up presenting functional explanations in the form of box and arrow diagrams which impute, but do not describe, causal sequences. A few psychologists continue to consider classical conditioning to be a psychological mechanism. I say a few psychologists because the cognitive revolution (Bandura, 1978; Dember, 1974; Gardner, 1985; Mahoney, 1977) arose by and large because most psychologists did not accept that cognition and other higher order cognitive processes could be explained by classical conditioning or, for that matter, by operant conditioning either. We also use imaging methods to identify brain structures that subserve psychological functions but do not, because we cannot, explain how these neural networks produce their psychological and/or behavioral effects (Dobbs, 2005; Uttal, 2001).
Gigerenzer (1998, 2009) emphasized the need to specify underlying mechanisms. Kazdin (2007) defined mechanism as “the basis for the effect, that is, the processes or events that are responsible for the change; the reasons why change occurred or how change came about” (p. 3). Squire, Knowlton, and Musen (1993) said, “Ultimately, one wants to understand cognition not just as an abstraction, or in terms that are simply plausible or internally consistent. Rather, one wants to know as specifically and concretely as possible how the job is actually done.” (p. 454). Kazdin (2008) clarified his meaning of mechanism as follows: “By mechanisms, I refer to the processes that explain why therapy works or how it produces change” (p. 151). A mechanism, therefore, consists of a sequence of causal events that are either necessary or sufficient to bring about the imputed result.
A search of Psych Info for “psychological mechanisms” on 11/11/11 returned 1,559 references, which demonstrates considerable interest in understanding how psychological phenomena are implemented. Unfortunately, most of these references fail to provide the details requested above. Instead, they impute causation without describing the sequence of causal steps that constitute causation; that is, detail how the mechanism operates. We currently need but do not have adequate mechanism information for most psychological phenomena. This includes how and why our empirically supported treatments work. We have yet to identify their active ingredients.
Proposed Connectionist Network Learning Theory
A potential solution to the above stated explanatory problems is presently available but has not yet been published in an organized way because investigators have not fully appreciated the broader implications of their work.
The origin of all that follows is the result of approximately 25 years of my reading and thinking about connectionist neural networks and their theoretical, empirical, and practical applications to psychological science and clinical psychology. When I first encountered connectionist network theory, my interest was restricted to exploring its potential applications to clinical psychology because I am a clinical psychologist. What I discovered is that this theoretical orientation has a remarkably diverse explanatory scope that applies to all or most areas of psychological science.
I begin by describing a generic connectionist neural network model introduced by McClelland and Rumelhart (1986) and Rumelhart and McClelland (1986) to define basic vocabulary. Then I discuss the explanatory nucleus, or kernel, of the theory that I introduced above. The three neural network properties in this nucleus necessarily occur during every processing cycle. They always occur together and cannot be separated. Alternatively stated, transformation and experience-dependent plasticity occur during, as an integral part of, and as a result of every network cascade. I present these three network properties as core psychological principles because they constitute an explanatory nucleus that is central to all psychological explanations that I provide.
Next I present eight corollary psychological principles that derive from this explanatory nucleus; that is, they can be explained using combinations of the three core principles in various proportions. It is important to emphasize the derivative nature of these eight corollary principles. They are true because the three core principles are true. It is equally important to emphasize that these principles are not new. Experimental psychologists have independently discovered them. No further empirical support is needed to warrant our acceptance of them. Finally, it is important to emphasize that having grounded this vocabulary and set of core and corollary principles in neuroscience helps integrate and unify psychology and biology; that is, makes psychology more consilient (Wilson, 1998). This approach is sometimes called connectionism by psychologists and computational neuroscience by biologists.
The explanatory nucleus of the theoretical approach advocated here consists of three elemental neural network properties: network cascade, transformation, and learning/memory mediated by experience-dependent plasticity (EDP). These three explanatory elements always occur together. Neural networks typically consist of three or more layers of processing nodes/neurons interconnected with two or more layers of connections/simulated synapses. When one or more node/neuron is activated in response to stimulation, that activation automatically and mechanically spreads across all connected network layers. Just as activation automatically spreads down the axon of a neuron, here we can envision activation spreading across an entire neural network. I refer to this process as the network cascade. It is the first and most fundamental explanatory element; the origin of every explanation that follows hereafter. This network cascade transforms sensory stimuli into cognitive or cognitive-like constructs in a process I refer to as transformation which is the third elemental network property. The aforementioned cascade also activates a complex biological mechanism known as experience-dependent plasticity that modifies the network; that is, changes network connections and modifies processing nodes. We experience these alterations as memory formation and learning. The modified network will now process differently. This constitutes the second elemental network property: learning and memory produced by EDP.
At least eight corollary psychological phenomena arise from this explanatory nucleus through a process known as emergence. These corollary psychological principles can then be used to explain a broad range of additional psychological phenomena. The resulting explanatory diversity is so large that these phenomena would not normally be discussed together, may appear to be unrelated, and may even appear to be random and disjointed and to lack adequate transition from one to the other in this article. But I aim to show that all these seemingly unrelated phenomena arise from the aforementioned set of core concepts and can be explained using a common vocabulary. Such explanatory diversity is remarkable within psychological science and demonstrates the relevance of the approach advocated here to a truly general psychology. This explanatory approach should appeal to psychologists who have quite different interests, and consequently should be understood as a strength of this theoretical orientation. It is worth emphasizing that I did not select the phenomena that I will be discussing in this article. Instead, these diverse topics emerged as I explored applications of connectionist network learning theory to psychology.
An additional advantage of the network approach presented here is that the associated computational models provide some of the missing mechanism information referred to above as the third explanatory problem. Processing by artificial neural networks can be mathematically specified in sufficient detail that the model can be simulated on a computer. These computational models enable investigators to perform true experiments that support cause and effect conclusions, not merely associations, because these models allow investigators to create and manipulate neural architectures, provide them with specific learning histories, and damage them in controlled and specific ways. Investigators can thus examine developmental processes in ways that might otherwise be impractical or unethical. Theoretical synthesis is also possible, as when Read et al. (2010) combined personality structure and dynamics into a single model.
Scientific versus Clinical Principles
Castonguay and Beutler (2006) used the term principle to refer to facts capable of guiding clinical practice, such as “Age is a negative predictor of a patient's response to general psychotherapy” (Beutler, Blatt, Alimohamed, Levy, & Angtuaco, 2006, p. 27), and “The more impaired or severe and disruptive the problem, the fewer benefits are noted for time-limited treatments” (Beutler, Castonguay, & Follette, 2006, p. 112). I counted 216 of these statements across 17 chapters. While useful in guiding clinical practice, these principles carry little if any explanatory power. The principles presented in this article are far fewer and are much more powerful because they can be used to explain a wide variety of psychological phenomena including why effective therapies work. The analysis I present in this article is not necessarily complete. There may be more than eight corollary principles that arise from the three core network properties that form the explanatory nucleus introduced above.
A Generic Connectionist Neural Network (CNN) Model
CNN models are mathematical representations of artificial and/or real neural networks. They are admittedly simplistic first approximations to real brains but provide a disciplined and detailed way to discuss complex psychological and behavioral phenomena. I will describe how they can be seen to process cognition and affect, resulting in simulated behavior; but, like astrophysical models of supernova that do not actually explode, these models do not actually think, feel, and behave. I introduce a generic CNN model here in order to clarify basic terms and concepts. The central and essential insight provided in the two volume seminal work by Rumelhart and McClelland (1986) and McClelland and Rumelhart (1986) is that one can effectively simulate many psychological phenomena using brain-inspired parallel-distributed processing (PDP) connectionist neural network models. This extraordinary achievement integrates biology (neuroscience) and psychology in a fundamental way by providing both sciences with a common vocabulary and a shared set of basic concepts. I aim to extend this integration by identifying what I consider to be core and corollary psychological principles that arise from CNN models.
Parallel distributed networks
One can begin to understand the ability of networks to do psychology by extending the once standard psychological perspective. Behavioral psychologists once theorized in terms of the Stimulus
Figure 1 illustrates a generic form of a simple CNN model. It mainly serves to illustrate a few fundamental features of these models and to provide a structure for describing basic terms and concepts. One can think of the S, O, and R elements as either neurons or more abstractly as psychological processing nodes. This model has two fundamental properties. First, it has a parallel form; that is, multiple copies of the S

Illustration of a double generalization of the S
Simulated dendritic summation
Real synapses are either excitatory or inhibitory to various degrees. Every connection, simulated synapse, between network nodes in a CNN model carries a numerical value that represents its degree of excitation or inhibition. These numerical values are called connection weights. They simulate synaptic properties of excitation through inhibition. Every line in Figure 1 carries such a weight (not shown). The sum of the weights associated with converging lines that enter nodes in the middle and lower layers is referred to as the Net Input to those nodes and simulates dendritic summation. In real neural networks, neurons can receive up to 10,000 inputs from other neurons.
Representation
Computer-based network models are mathematical simulations implemented in software and, therefore, do not actually perceive, think, or behave; any more than mathematical astrophysical models of supernova explode. However, there are hardware simulations of neural networks such as those by Kwabena Boahen, which are based on neuromorphic chips that use the same physical forces to move electrons through transistors as cells use to pass ions across cell membranes. These models offer the potential for more psychologically realistic simulations (http://www.youtube.com/watch?v=mC7Q-ix_0Po, http://www.stanford.edu/group/brainsinsilicon/). Good models, however, reflect important features of the phenomena they simulate. Most connectionist models simulate only a few basic synaptic properties and yet do surprisingly well at simulating psychological phenomena. The models created by computational neuroscientists differ from the connectionist models created by psychologists in their goals and objectives; they focus on simulating basic neuroscience phenomena rather than psychological phenomena.
Stimulus or S-layer
Because connectionist models cannot actually perceive stimuli, the models must represent basic stimulus features. One way is to code for relevant stimulus components called microfeatures. For example, representation of the perception of a cup of coffee includes the following stimulus microfeatures: type of container (e.g., paper cup, mug, or traditional cup), container size (e.g., large, medium, or small), amount of cream (e.g., none, some, more, or a lot), amount of sugar (e.g., none, some, more, or a lot), and temperature (e.g., cold, hot, or very hot). In this model, the No. 1 represents the presence of a microfeature; 0 represents its absence. A thermometer code can be used to represent quantitative variables. For example, a 1–7 scale for the intensity or degree of some microfeature can be represented by seven S-elements where 0000000 indicates complete absence of a microfeature, 0000001 indicates very little, 0000111 indicates more, and 1111111 indicates the maximum amount of a particular microfeature.
A second way to represent stimulus microfeatures is to code what an artificial retina might see. For example, one can accomplish the perception of the letter “A” using the scheme illustrated in Figure 2 where each of 5 blocks (rows) of 5 microfeatures combine to represent the letter “A” on an artificial retina. The first block of five digits (0, 0, 1, 0, 0) corresponds to the off = 0 = white and on = 1 = black pattern found in the top row of Figure 2. The second through fifth blocks of digits code for the off-on pattern of the “A” in rows two through five of Figure 2. Hence, the single row of 25 microfeatures below the boxes represents “seeing” the letter “A.”

Illustration of how perceiving (S-layer) and saying (R-layer) the letter “A” results in turning 25 stimulus microfeatures on or off. The first block of five digits correspond to the off = 0, on = 1 status of the five pixels in the top row of an artificial retina. The remaining blocks of five digits describe the on-off status of the pixels in the remaining rows.
Transformation or O-Layer
The nodes in the O-layer of Figure 1 represent cognitions and emotions, that is, psychological features generated by the organism by transforming stimulus microfeatures. I describe the transformation process in greater detail later in this article (as Principle 3 below).
Response or R-layer
The last layer of the network codes for behavior. The simplest response layer contains a single node which is either on or off indicating presence or absence of a single behavior. For example, behavior in a Skinner box can be represented with a single R-node that is either on meaning that the bar is pressed = 1 or off meaning that the bar is not pressed = 0. If the Skinner box has two levers, that is, the concurrent operant case, and the animal must choose which lever to press, then two R-nodes are needed to represent both possible behaviors. Representing more complex behaviors requires more R-nodes. The scheme used to represent perception (see Figure 2) can also be used to represent behavior.
The Explanatory Nucleus: Three Core Psychological Principles
Principle 1: The Network Cascade; Unconscious Processing
The activation cascade across multiple neural network layers constitutes the origin and genesis of all explanations provided from the perspective advocated here. This cascade process occurs unconsciously, automatically, and mechanically, like neural conduction down an axon except now we have parallel distributed processing across a web consisting of layers of synapses (connections) among neurons (processing nodes). Two causal consequences of this network cascade always occur during and as a result of every cascade processing cycle. They are Experience-dependent plasticity mediated learning and memory formation, discussed as Principle 2, and Transformation, discussed as Principle 3.
The role of unconscious processing has long separated psychodynamic psychologists from their cognitive–behavioral colleagues and constitutes a serious schism in psychology (Staats, 1983). There is nothing about CNN models that requires consciousness; the cascade process occurs automatically and mechanically, much like an action potential down a neuron, so consciousness need not be assumed. Kihlstrom, Barnhardt, and Tataryn (1992) noted that the cognitive unconscious is not as extreme as Freud conjectured. The following well-accepted findings in psychological science further document the existence of unconscious processing.
Cognitive science
It is very well documented that people automatically, and therefore unconsciously, use cognitive short cuts called heuristics that result in well-documented biased conclusions (cf. Kahneman, 2003; Tversky & Kahneman, 1974) called cognitive illusions (Kahneman, 2011, p. 27).
Cognitive neuroscience
Over a half century of research on the split-brain (Gazzaniga, 2005) has clearly revealed unconscious processing by our left and right hemispheres.
Social psychology
Social psychologists have empirically demonstrated that automatic unconscious processes causally influence social cognition (Bargh, 2007; Gawronski, & Payne, 2010) including social cognition of close relationships such as transference and counter transference (Chen, Fitzsimons, & Andersen, 2007). Social psychology now has an Unconscious-thought Theory (UTT; Bargh & Morsalla, 2008; Dijksterhuis & Nordgren, 2006) supported by empirical evidence that unconscious processing can be superior to conscious processing when making complex decisions.
Theoretical reorientation
Contemporary cognitive psychology assumes that people are conscious. Consequently, Bargh and Morsalla (2008) have characterized contemporary cognitive psychology as conscious-centric. But because conscious-centric psychology assumes consciousness, it cannot explain it because one reasons from, not to, axioms. Explanations of how consciousness arises therefore lay outside the explanatory scope of conscious-centric psychology.
Presuming rather than explaining consciousness makes it difficult to determine where consciousness processing ends and unconscious processing begins. This problem has impeded research into unconscious processing from a conscious-centric perspective for decades (Erdelyi, 1992). Conscious-centric psychologists therefore either deny that unconscious processing exists or claim that it lies outside the relevant explanatory sphere of their science. Conversely, the unconscious-centric position need only show how consciousness emerges. The network perspective introduced here does not assume consciousness and, therefore, is not precluded from explaining it. Consequently, the unconscious-centric perspective constitutes a more fundamental theoretical position. Replacing our current conscious centric orientation with an unconscious-centric one is perhaps the most important contribution made by the network perspective advocated here; it constitutes a major paradigm shift for psychological science (cf., Kuhn, 1996).
Principle 2: Experience-Dependent Plasticity Mediates Learning and Memory
An immediate unavoidable consequence of the automatic unconscious CNN cascade is that it changes the network through a process known to neuroscientists as experience-dependent plasticity, which is the biological basis of learning and memory.
The following material provides specific mechanism information for how learning and memory work that is included in CNN models in one form or another. At one time psychologists studied learning as a basic process (Bower & Hilgard, 1981). But psychologists largely abandoned this study when they adopted an information processing orientation. Meanwhile, neuroscientists identified the biological mechanisms of learning and memory discussed below and more completely by Bear, Connors, and Paradiso (2007), Carlson (2010), Hell and Ehlers (2008), Kalat (2009), Lisman and Hell (2008), and Squire et al. (2008). This biological mechanism information explains previously unanswered basic psychological science questions such as why reinforcers are reinforcing and/or why reinforcers modify behavior. Cognitive psychologists largely ignored operant conditioning because its simple R
Carlson, Miller, Heth, Donahoe, and Martin (2010) defined learning in terms of memory, “Learning refers to the process by which experiences change our nervous system and hence our behavior. We refer to these changes as memories” (italics in the original; p. 440). Alternatively stated, learning and memory represent two sides of one coin. Gluck and Myers (1997) referred to computational (CNN) models of learning and memory as providing “… the conceptual glue to bind together data from multiple levels of analysis” (p. 481, italics added) including “cellular, physiological, anatomical, and behavioral levels” (p. 510).
Experience-dependent brain/network plasticity
Whereas the computers that implement connectionist models remain unchanged by the processing they do, the neural networks they simulate are changed by the processing they do (Martin, Grimwood, & Morris, 2000). This simulates the fact that brains have a dynamic and adaptive quality that computers currently lack. Moreover, learning and memory occur within the same networks, implying that learning modifies memory and, therefore, that memory changes can track learning (Tryon & McKay, 2009). Experience-dependent plasticity is the term used to refer to the ways in which the brain changes during learning and memory formation. Neurons that fire together wire together through the synthesis of new proteins resulting in Long Term Potentiation (LTP) and Long Term Depression (LTD) of synapses depending upon details that lie beyond the scope of this article (cf., Bear, Connors, & Paradiso, 2007; Carlson, 2010; Hell & Ehlers, 2008; Kalat, 2009; Lisman & Hell, 2008; Squire et al., 2008). LTP enables incoming signals to produce a stronger response. Evidence exists for the following three mechanisms: (a) Presynaptic terminals may release more neurotransmitter, (b) the number of postsynaptic receptors may increase, and/or (c) present receptors may become more responsive to neurotransmitter (Lombroso & Ogren, 2008, 2009). Structural changes may also occur. Dendritic spines may increase in length (Bear et al., 2007) and/or other spines may grow at the same location (Lombroso & Ogren, 2008, 2009). Neurons that fire out of sequence strongly reduce the flow of calcium ions through NMDA receptors and create the LTD inhibitory state that produces opposite physical changes from those of LTP: (a) Presynaptic terminals may release less neurotransmitter, (b) the number of postsynaptic receptors may decrease, and/or (c) present receptors may become less responsive to neurotransmitter. These and related dendritic spine modifications enable the brain to adapt to an ever-changing world throughout the life span, albeit in a gradually decreasing way. Together, these biological processes modify and literally sculpt, brain networks thereby adapting them to their current physical and social environments. This process enables life span development to occur.
The experience-dependent plasticity mechanism is routinely incorporated into connectionist models of many psychological phenomena (Abbott, 2008; Arbib, 2002; McClelland & Rumelhart, 1986; McLeod et al., 1998; O'Reilly & Munakata, 2000; Rumelhart & McClelland, 1986; Sun, 2008). The unconscious CNN cascade that automatically transforms stimulus microfeatures into latent constructs and adaptively changes the network through experience-dependent plasticity is a phylogenetically general learning/memory mechanism that constitutes a modern cognitive-affective neuroscience learning theory. I use the word theory to emphasize the explanatory value of this learning mechanism. Unfortunately the term theory denotes tentative acceptance pending subsequent empirical validation. Nothing could be further from the truth here, as this understanding of learning and memory is based on neuroscience facts that have been empirically established by any reasonable standard.
Principle 3: Transformation
Another unavoidable consequence of the automatic unconscious network cascade process that occurs during, and as a result of, the network cascade, is that stimulus microfeatures get transformed into cognitions, expectations, or cognitive-like constructs. How brain networks can possibly form psychological concepts has long been a mystery. CNN models offer a possible explanation.
We form concepts by processing stimulus microfeatures. For example, the stimulus microfeatures of wings, feathers, and beaks are positive indicators, and defining features, of the concept “bird.” Antlers are negative indicators of the concept “bird.” If it has antlers it cannot be a bird. In factor analytic terminology we say that wings, feathers, and beaks load positively and antlers load negatively on the latent construct/concept “bird.” CNN models also use factor analysis to form constructs from microfeatures. 1 Psychologists have accepted factor analysis as valid means of identifying latent constructs from psychological test items since Spearman (1904) introduced this methodology over a century ago.
The top half of Figure 3 illustrates a typical two factor solution of six test items. It is important to emphasize that the computer generates the latent factor constructs and that psychologists merely name them. The bottom half of this figure represents the two factor model as a CNN model by inverting (vertically flipping) the factor analytic model presented in the top half of Figure 3. After flipping, the top row of six S-nodes (V1−V6) can now be understood to represent six stimulus microfeatures. The second row of two O-nodes (F1 & F2) continues to represent two computer-generated latent constructs/concepts.

The top half of this figure illustrates a two factor (F1 & F2) model of responses to six psychological test items, variables (V1−V6). The computer calculates factor loadings, illustrated generically with lines that define the two latent constructs and show the importance of each item to each factor. The investigator names the constructs that the computer constructed. The bottom half of this diagram vertically inverts the two factor diagram into a two layer network model where the top row of six S nodes (V1−V6) now represent six stimulus microfeatures and the second row of two O nodes (F1 & F2) continue to represent two latent constructs that are transformations of the stimulus microfeatures.
Further insight into how neural networks construct constructs can be obtained from Loehlin's (1987) description of Cyril Burt's (1917) centroid method of factor extraction. What follows is an effort to provide some computational mechanism information that entails high school algebra. Table 1 presumes that the connection weights, path coefficients, in Figure 3 are labeled a through f from left to right, respectively. Table entries consist of all possible pairwise paths.
Centroid Factor Extraction Based on Loehlin's (1987) Description of Cyril Burt's (1917) Methods. Entries Consist of All Possible Pairwise Paths
Let
Then the sum associated with the Path a column equals aΣ, the sum associated the Path b column equals bΣ, and so forth through the sum of the Path f column equals fΣ. The sum of the column sums, the grand sum, equals
Redefining
Summation is the main operation here and neurons are very good at dendritic summation. While it is less clear how real neural networks might accomplish division and square root extraction, simulated networks have no trouble doing so. The end result is network factor extraction that entails factor extraction methods that psychologists have used for more than a century to discover latent constructs. Application of these methods by artificial neural network models authorizes one to call the extracted factors simulated constructs or concepts. Psychologists who have accepted the legitimacy of factor analytic concepts for the past century must now also accept the legitimacy of CNN derived concepts. In any case, the O-nodes compute a highly transformed version of the S-node microfeatures. The R-nodes compute second-order factors.
The ability to at least twice transform stimulus microfeatures qualitatively distinguishes CNN models from behavioral models and certifies them as cognitive models. CNN models of classical conditioning therefore authorize discussions of Pavlovian conditioning in terms of expectations (cf., Rescorla, 1988). It is noteworthy that the Rescorla-Wagner model of classical conditioning (Rescorla & Wagner, 1972) is mathematically equivalent to the Delta learning function used in some connectionist models.
Corollary Psychological Principles
Corollaries are true because what they are corollary to is true, as in mathematics where corollaries follow from theorems. The psychological principles presented in this section necessarily follow from the explanatory nucleus of the three core principles presented above. This means that these corollary principles are also intrinsic, inherent, essential, and basic, properties of connectionist networks.
Principle 4: Priming
The APA concise dictionary of psychology (APA, 2009) defines priming in cognitive psychology as “… the effect in which recent experience of a stimulus facilitates or inhibits later processing of the same or similar stimulus” (p. 395). Priming is a consequence of the CNN cascade plus learning and memory mediated by experience-dependent plasticity. These core principles explain why priming phenomena exist and how they work. Consider repetition priming. Subliminal or initial stimulation automatically and unconsciously spreads, that is, cascades, across the network. This processing activates experience-dependent plasticity mechanisms that strengthen the particular processing pathways taken across the network. These changes increase the likelihood that subsequent subliminal and/or supraliminal stimuli will follow these processing pathways. Alternatively stated, the processing pathways taken by subliminal or initial activations are biologically reinforced, strengthened by the experience-dependent plasticity process, thereby increasing the probability that subsequent activations will follow the same processing pathways. Strong empirical support exists for repetition priming, semantic (lexical) priming, conceptual priming, emotional priming, cultural priming, and idea-motor priming.
Principle 5: Part-Whole Pattern Completion
Part-whole pattern completion is a consequence of the CNN cascade and transformation. It is known to psychologists by several other names. The general term, content addressable memory, refers to the fact that some content can retrieve memory for related content; that is, one fact can cause a person to recall other related facts. Mood-congruent recall, also known as state dependent memory, is a more specific form of content addressable memory that derives from the fact that emotions are encoded along with cognitions and consequently, constitute partial cues. People recall more positive events when happy than sad and more negative events when sad than happy (cf., Blaney, 1986; Bower & Cohen, 1982; Isen, 1984; Matt, Vazquez, & Campbell, 1992; Teasdale & Fogarty, 1979; Williams, Watts, MacLeod, & Mathews, 1988). Memory retrieval is enhanced when retrieval mood matches encoding mood (Eich & Macauley, 2006). Other examples of this principle come from perception, PTSD (Tryon, 1999), forensics, false confessions, (Brainerd & Reyna, 2005), and imagination inflation (Garry & Polaschek, 2000).
Principle 6: Consonance
Mathematical properties of the equations used to simulate network cascade activated experience-dependent plasticity cause CNN models to seek consonance, that is, to generate increasingly coherent results. Heider (1958) noted that people seek consonance, also known as coherence. Festinger (1957, 1964) rephrased this motive as avoiding dissonance. These consistency theories have received extensive empirical support and constitute of some of the best replicated findings that psychological science has to offer (cf., Abelson et al., 1968; Read & Miller, 1998). The consonance seeking property of CNN models has enabled them to accurately simulate the results of consonance/dissonance psychological experiments.
Consonance seeking is also emotionally motivated. Kunda's (1990) seminal paper entitled, “The Case for Motivated Reasoning” made the case for “hot cognition” as follows:
There is considerable evidence that people are more likely to arrive at conclusions that they want to arrive at, but their ability to do so is constrained by their ability to construct seemingly reasonable justifications for these conclusions (p. 480). People do not seem to be at liberty to conclude whatever they want to conclude merely because they want to. Rather, I propose that people motivated to arrive at a particular conclusion attempt to be rational and to construct a justification of their desired conclusion that would persuade a dispassionate observer. They draw the desired conclusion only if they can muster up the evidence necessary to support it. In other words, they maintain an ‘illusion of objectivity’… To this end, they search memory for those beliefs and rules that could support their desired conclusion. They may also creatively combine accessed knowledge to construct new beliefs that could logically support the desired conclusion. It is this process of memory search and belief construction that is biased by directional goals (pp. 482–483).
That people unconsciously and automatically selectively attend to and “cherry pick” information from memory, archival sources, and their current environment to support previously arrived at emotionally held convictions while maintaining the illusion that they are being objective is very familiar to social and clinical psychologists.
Principle 7: Dissonance Induction/Reduction
Students and clinicians are faced with more treatments than they can learn or apply and little basis to choose. We have yet to identify the active ingredients that explain why diverse treatments work. The Dissonance Induction/Reduction Principle extends the Consonance Principle discussed above and was formulated to explain why treatments work (Tryon, 2005; Tryon & Misurell, 2008). It is based on how CNN models are trained. Yes, investigators train rather than program CNN models. Training begins by presenting a stimulus that activates the relevant microfeature indicators in the S-layer. In the simple generic network illustrated in Figure 1, processing automatically spreads across the first layer of connections to the processing nodes in the O-layer and automatically continues across the second layer of connections to the processing nodes in the R-layer where it activates some R-nodes and not others, and/or the processing deactivates some R-nodes. The result in either case is a string of 1's and 0's indicating which R-nodes are on and which R-nodes are off. The investigator compares this digital string to the desired one. The number of 1's that should be 0's and the number of 0's that should be 1's defines the Hamming distance between the computed and desired response and is a measure of dissonance. Investigators use a learning function such as the Back Propagation algorithm to reduce these discrepancies. The Back Propagation method first changes the connection weights between the R- and O-nodes and then changes the connection weights between the O- and S-nodes; that is, works backward through the network relative to the normal processing flow. Mathematical properties of the Back Propagation algorithm used to simulate experience-dependent plasticity insure that the Hamming distance becomes smaller with each training trial thus reducing dissonance. In this way, the network learns to make the desired response when presented with specific stimuli through a process of dissonance reduction. I now use this training method to explain how empirically supported treatments (ESTs) work.
EST mechanism information
Tryon (2005) reviewed possible mechanisms for why systematic desensitization and exposure therapy work and found all of them defective. He then presented a connectionist mechanism that Tryon and Misurell (2008) extended to depression and formulated as a Dissonance Induction/Reduction Principle. This principle stems directly from how connectionist network models are trained described above. It also explains why the unified protocol described by Allen, McHugh, and Barlow (2008) for treating emotional disorders, including anxiety and depression, based on over 25 years of clinical research, works. It also explains why motivational interviewing, a treatment that explicitly uses dissonance reduction therapeutically, works (Miller & Rollnick, 2002; Tashiro, & Mortensen, 2006).
Prior to treatment, fear stimulus microfeatures represented across the client's S-nodes activate cognitions and emotions represented across the O-nodes, which then activate avoidance behaviors represented across the R-nodes. Thus, the client avoids or escapes from the situation. The therapist induces dissonance by arranging for new responses to occur. This dissonance-inducing procedure is equivalent to “clamping” the R-nodes to different values than the ones that the network computed prior to treatment. This increases the Hamming distance, a dissonance measure, between the pretreatment response of avoidance and the desired therapeutic response. We established above that people, and connectionist networks, naturally seek coherence. Hence, the client's subsequent processing reduces dissonance and increases consonance. After repeated trials, processing has modified the connection weights such that the behavioral result computed by the network is more consistent with the desired behavior prescribed by the therapist than was the initial behavioral response. Alternatively stated, before treatment, the client processes one or more anxiety related stimulus microfeatures into fearful thoughts, feelings, and avoidance responses. Dissonance results when a prescribed therapeutic response differs from the client's pretreatment response. Dissonance reduction occurs when the client remains in the situation until his or her network settles into a more congruent state. The client now thinks and feels differently because their network has been physically modified by experience-dependent plasticity mechanisms in ways that are more congruent with therapist prescribed behaviors. This understanding can be communicated to clients to help them understand the necessity to carefully carry out prescribed treatments. Therapists should look for, and maximize, this active ingredient in any treatment they choose to use.
Principle 8: Memory Superposition
Unlike computers that place memories in unique physical locations, people and CNN models store multiple memories in the same neural networks. Bechtel and Abrahamsen (1991, pp. 70–81) presented a simple numerical example to illustrate this process. CNN networks can store and retrieve more memories together more effectively and accurately if their content is distinct. Similar to people, CNN models tend to confuse memories that share common elements, which makes them good psychological models. CNN models use the same networks both for processing and storage. Processing changes the network connections via experience-dependent plasticity mechanisms and therefore learning-based treatments can modify memory such that memory changes can be used to document treatment effects (Tryon & McKay, 2009).
Principle 9: Prototype Formation
Prototype formation is another consequence of the CNN cascade and the transformation principle. The APA Concise Dictionary of Psychology (2009) defines a prototype as “in the formation of concepts, the best or average exemplar of a category. For example, the prototypical bird is some kind of mental average of all the different kinds of birds of which a person has knowledge or with which a person has experience” (APA, p. 402). Exemplars of a prototype are stimuli that possess some of the defining features of the prototype. Exemplars can vary from exhibiting just one defining feature up to exhibiting one less than all of the defining features.
CNN models form prototypes by extracting averages from repeated presentations of stimuli over time as part of the transformation property of the network cascade. Each stimulus represents the prototype plus deviations. The prototype gets reinforced over trials whereas the deviations vary and tend to cancel one another. Experience-dependent plasticity mechanisms biologically reinforce our brain networks response to the common features and thus, results in prototype formation.
Prototypes are important to cognitive psychology because they facilitate, that is, speed up, the processing that humans do. Winkielman, Halberstadt, Fazendeiro, and Catty (2006) presented evidence to demonstrate that prototypes facilitate emotional and cognitive processing. Prototypes function like stereotypes (e.g., Cin et al., 2009). Therefore, being able to simulate prototype formation is important.
Principle 10: Graceful Degradation (Aging)
Understanding the effects of age on cognitive performance is increasingly important for psychologists. CNN models provide a succinct way of understanding many cognitive effects of aging. In a remarkable series of simulations, Li, Lindenberger, and Sikström (2001), Li and Sikström (2002), and Li, Naveh-Benjamin, and Lindenberger (2005) successfully modeled several important features of aging by changing a single network parameter that changes the slope of the sigmoidal transfer function that determines the probability of firing from Net Input. Networks with steeper transfer functions performed like younger people on several experimental tasks. Networks performed like older people as the slope of their transfer function decreased. Investigators effectively simulated the behavior of young, middle aged, and old people by using three different slope levels. The results of these computer simulations agreed remarkably well with data from experiments with real people. Another aspect of CNN models that implement parallel-distributed-processing (PDP) is that their performance degrades gracefully, slowly, given neurological insults of various kinds just like people do during normal aging.
Principle 11: Top-Down & Bottom-Up Processing
Psychologists have long debated whether emotions have their origins in perception (e.g., LeDoux, 2000; Phelps, 2006) or in cognition as Beck (1972), Ellis (1994), and others have argued. A network perspective renders this theoretical schism a false dichotomy. Ochsner et al. (2009) used fMRI and found that emotion generation involves both top-down and bottom-up processes. Network processing flows in two directions. Bottom-up processing begins in sensory organs, such as the eye and ear, and terminates in the visual and auditory cortices. Top-down processing begins in higher brain centers as expectations that influence perception and are constrained by attention. Freeman and Ambady (2011) presented a connectionist neural network simulation of person construal based on the continuous interaction between top-down processing of high-level cognitive states (regarding social categories and stereotypes) and low-level bottom-up processing of facial, vocal, and bodily cues.
Placebos and nocebos are examples of top-down processing based on suggestion. They can be explained as the result of priming, which was explained here as a network property that can be understood in terms of the explanatory nucleus and its three core principles introduced above. Hallucinations are pathological examples of how expectations change perception. Frank, Cohen, and Sanfey (2009) characterized the bottom-up processing system as automatic, intuitive, emotional, and implicit and the top-down system as controlled, deliberative, and explicit. They documented that both systems operate simultaneously and recommended computational models to better characterize how these two systems interact. This principle emphasizes that the cascade process simultaneously flows in opposite directions.
Salzman and Fusi (2010) presented evidence that the anatomical interconnections between the amygdala and the prefrontal cortex result in the simultaneous operation of these systems, thus seriously undercutting the CBT theoretical premise that cognition causes emotion and not the reverse. The ability of network models to inform psychologists regarding how cognition and emotion interact and reciprocally influence one another is a major contribution to the science and practice of psychology. The network perspective replaces simplistic either/or views on the generation of emotions with a more comprehensive understanding that is consistent with neuroscience evidence.
Review and Summary
Table 2 reviews and summarizes the three core principles that constitute the explanatory nucleus, the eight corollary principles, and provides examples of additional psychological phenomena that they can explain.
Core & Corollary Principles
These three core network properties function in the same way as Piagetian assimilation/accommodation cycles do.
Criticisms and Rebuttal
Some authors have criticized the entire effort to integrate psychology and biology as a form of reductionism. Other authors have criticized CNN models as little more than curve fitting (Gigerenzer, 2009). Next, I address both concerns.
Reductionism
Reductionism is the simpler of two possible relationships between psychology and biology. Reductionism seeks to break down psychological phenomena into constituent biological elements that can be used to fully understand psychological phenomena. Efforts to reduce psychological states and constructs to physiological ones have failed and may never succeed for reasons presented by Miller and Keller (2000). The main problem with reductionism is that identification of biological elements such as brain structures and neurotransmitters does not explain how they interact to produce psychology and behavior. One can construct a biological explanation from these elements but not a psychological explanation. Connectionism endeavors to show how psychological phenomena can emerge from a network understanding of the interaction of basic neuroscience facts. The deterministic component of connectionist models is frequently mistaken for reductionism.
Emergence
Emergence concerns the more complex of two possible relationships between psychology and biology. The APA Concise Dictionary of Psychology (2009) defines emergence as “the idea that complex phenomena (e.g., conscious experiences) are derived from arrangements of or interactions among component phenomena (e.g., brain processes) but exhibit characteristics not predictable from those component phenomena” (APA, p. 163). For example, the physical properties of liquid water differ qualitatively and dramatically from those of its gaseous constituent elements of hydrogen and oxygen. The emergent properties of water derive from its physical structure; that is, how its atoms are interconnected. 2 CNN models presume that neurobiological processes synthesize all psychological phenomena through a complex emergent process. More specifically, the psychological and behavioral properties of real neural networks emerge from their genetic and epigenetic controlled architecture and the pattern of experience-dependent determined excitation and inhibition across all connections among processing nodes. These emergent network properties differ qualitatively from their constituent components as much as water differs from its constituent elements of hydrogen and oxygen gases. CNN models provide a unique opportunity to systematically, empirically study the process of emergence because investigators have complete control over every facet of network development and can repeatedly pause the emergent process to learn more about it.
Emergence of psychology from biology
Functional MRI (fMRI), and other brain imaging technologies enable investigators to observe and quantify brain activation occurring during psychological tasks. Cognitive neuroscientists have focused almost exclusively on identifying brain areas that mediate various psychological functions. 3 These methods support biological rather than psychological explanations. Psychologists want to know how neural networks give rise to cognition. Our discussion of how the automatic CNN cascade transforms stimulus microfeatures into latent constructs and activates experience-dependent plasticity mechanisms that mediate learning and memory, which arguably form the basis of all psychology, conceptually integrates psychology and neuroscience.
Falsifiability
At the highest level of abstraction, CNN models as a conceptual class are not falsifiable because, as Gigerenzer (2009) correctly noted, they can contain so many free parameters that they can fit any data. However, at the individual level, specific CNN models regularly get constrained, revised, and rejected on the basis of their ability to make specific predictions that are consistent with empirical facts; and that renders them both falsifiable and scientific (cf., Abbott, 2008). Some connectionist models are constrained by more neuroscience facts than others are.
Simulations
Parallel distributed processing connectionist neural network models develop their “adult” properties through many complex interactions with a simulated environment. These causal sequences are sufficiently complex that they cannot be adequately characterized by verbal descriptions alone. These sequences must be simulated using digital computers for now, and with physical models in the future, in order to determine the end result of many learning trials with sufficient precision to render the model falsifiable. Simulations can always be criticized for being artificial, incomplete, partial, and therefore, inaccurate implementations of the phenomena that they represent. This is also true of all laboratory-based psychological studies. Some critics view these limitations as fatal and reject all controlled study including connectionist models. More moderate minds recognize that the complexity of modeling complex psychological phenomena requires computational simulations despite their limitations. The field of psychometrics has long recognized, and benefited from, computer simulations and mathematical models.
Connectionist simulations provide us with powerful new methods. (a) They enable us to examine complex developmental sequences one step at a time to better see how experience-dependent phenomena emerge. (b) They enable precise control over initial conditions and developmental histories that may not be possible with living organisms. (c) They enable us to damage the network in very controlled ways that are presently technically impossible, impractical, and/or unethical in order to explore brain-behavior relationships. (d) They can facilitate theoretical integration. For example, Read et al. (2010) integrated personality dynamics and structure into a single model that agrees well with empirical findings.
Mathematical Proof
Minsky and Papert (1969) mathematically demonstrated the inadequacy of two-layered (S
Mature Science
Mature sciences have a common vocabulary and set of core principles. The connectionist network perspective introduced here uses basic neuroscience vocabulary and network concepts to construct mathematical models that have been shown to effectively simulate a wide variety of psychological phenomena (e.g., Abbott, 2008; Arbib, 2002; McClelland & Rumelhart, 1986; McLeod et al., 1998; O'Reilly & Munakata, 2000; Rumelhart & McClelland, 1986; Sun, 2008). I submit that this network perspective constitutes a modern cognitive-affective neuroscience learning theory that can help transition psychology into a mature science.
Footnotes
1
Some CNN models use principle components analysis (PCA) factor analysis to form concepts as part of their training (e.g., McLeod et al., 1998; O'Reilly, & Munakata, 2000). Connectionist models are part of mathematical psychology (Bower, 1994). They are also intimately related to multivariate statistics (White, 1989).
