Abstract
In two experiments, participants were given extinction training in a human causal learning task. In both experiments, three critical experimental cues were paired with different outcomes in a first phase of training and were then extinguished in a second phase. Three control cues were given the same treatment in the first phase of training, but were not then presented in the second phase. Participants’ ability to correctly identify the outcome with which each cue had been paired in the first phase was lower for extinguished than for control cues. Causal attributions to the extinguished cues were also lower than those to the control cues, a difference that correlated with outcome memory. These data are consistent with the idea that extinction in causal judgement is due, at least in part, to a failure to remember the cue–outcome relationship encoded in the first phase of training.
In an extinction procedure in Pavlovian conditioning, a conditioned stimulus (CS) is paired with an unconditioned stimulus (US) in a first phase of training and is then presented in the absence of that US in a second phase. The consequence of presentations of the CS alone in the second phase is that the conditioned response (CR) acquired by the CS in the first phase of training declines. Extinction is a very robust phenomenon and is observed in both animals (Pavlov, 1927) and humans (Paredes-Olay & Rosas, 1999).
Over the past decade or so, a dominant model of extinction has emerged, according to which extinction represents a failure to retrieve from memory the CS–US relationship encoded in Phase 1. This retrieval failure occurs as a consequence of interference from the more recent CS–no-US presentations in Phase 2 (Bouton, 1993). Support for this view comes from the observation of retroactive interference in paired-associate learning in humans (see Slamecka & Ceraso, 1960, for a review). When two words are paired in a first phase (A–B), and then Word A is presented with a different word, C, in a second phase (A–C), accuracy in recall of the A–B pair declines. Thus, it seems reasonable to suppose that, in the same way, CS–no-US pairings in Phase 2 of extinction training will prevent retrieval of the CS–US pairing encoded in Phase 1.
The present paper tests whether the extinction observed in human causal learning is due to a memory failure as suggested by Paredes-Olay and Rosas (1999). Human causal learning has been claimed to be an analogue of Pavlovian conditioning and so to be due to the formation of mental links between the cue and its outcome (Dickinson, 2001). Thus, cause is attributed to a cue to the extent that it is able to activate the outcome representation in memory. The retrieval interference account of extinction (Bouton, 1993) is based on the same mechanism; it is argued that the link formed in Phase 1 is not effective on test (the cue does not activate the outcome) due to interference from Phase 2.
Despite the broad consensus among animal learning researchers that extinction is due to a retrieval failure, there are other explanations. Extinction might be observed, despite perfect memory retrieval from both phases, as a consequence of an integration of cue–outcome and cue–no-outcome information on test (see Pineno & Miller, 2005, for review). Thus, because the extinguished cue is only sometimes followed by the outcome, it is deemed not to be a strong cause of this outcome. This would be consistent with an inferential reasoning approach to human learning (e.g., De Houwer, Beckers, & Glautier, 2002; Lovibond, Been, Mitchell, Bouton, & Frohardt, 2003; Waldmann, 2000). The strongest support for this position comes from the phenomenon of instructed extinction; merely informing the participant that the cue no longer signals the outcome is enough to immediately reduce responding on test (Grings, Schell, & Carey, 1973).
Whether extinction in causal learning is due to a failure of memory can be tested by combining a standard causal learning task with a more direct measure of memory. Just such a technique has been used to examine the role of cue–outcome memory in another learning phenomenon, blocking (Mitchell, Lovibond, Minard, & Lavis, 2006). In blocking, little is learned about a target cue (and so low causal ratings are given on test) if it is trained in compound with another cue that already predicts the outcome. Mitchell et al. (2006) examined whether blocking in causal judgements was accompanied by a failure to remember which outcome had been paired with the blocked cue. They used an “allergist” task in which the fictitious patient, Mr X, consumed a range of foods, some of which were followed by an allergic reaction. A number of different allergic reactions were used, and the participants’ task on test was to judge the causal status of each food cue, and to remember the specific outcome with which that cue had been paired. Mitchell et al.'s (2006) participants not only gave the blocked target cues lower causal ratings, but they also failed to remember which outcome had followed those target cues in training. This lends support to the idea that blocking results from a failure to encode the cue–outcome pairing in memory.
In the current experiments, we paired a range of food cues with a variety of fictitious allergic reactions in a first training phase and extinguished some of those cues in a second phase. Participants were then asked to remember the outcome with which each cue was paired in the first phase and to judge the causal efficacy of each cue. The theory that extinction is a memory failure predicts two things. First, participants’ identification of the outcomes paired with the cues in the first training phase will be poorer for those cues that were later extinguished than for their nonextinguished controls. Secondly, extinction in memory should correlate with the extinction seen on the causal judgement measure. Thus, the present study tests whether extinction in causal learning is due to a failure of memory, as suggested by Paredes-Olay and Rosas (1999).
EXPERIMENT 1
The extinction and control cues (A and B, respectively) were paired with an allergic reaction outcome (O) during the first phase of training, and A–no-outcome trials were presented in the second phase. The design was implemented three times with three different fictitious allergic reaction outcomes. Cues A1–A3 were paired with O1–O3 in Phase 1 and were extinguished in Phase 2. Cues B1–B3 were paired with O1–O3 in Phase 1 but were not presented in Phase 2.
This experiment was conducted in two separate replications, which used slightly different test instructions. In both replications, participants were asked to indicate, on test, the outcome that had followed each cue during training. However, in the second replication, it was emphasized that the participants were to report the outcome that had followed the food “at any stage throughout the experiment”, in order to ensure that they identified the outcome paired with A in the first phase. This is discussed further following presentation of the method and results.
Method
Participants
Participants were 60 undergraduate first-year psychology students at the University of New South Wales, who completed the experiment in partial fulfilment of a course requirement. Of these, 31 participants completed the first replication, and 29 completed the second replication.
Apparatus and stimuli
All components of the task were presented, and all responses recorded, on a PC running a program written in Revolution Studio, Version 2.7.0. The cues were 15 foods (pear, cheese, lemon, yoghurt, peaches, rice, ham, mushrooms, coffee, garlic, steak, margarine, chicken, banana, and tomato). The outcomes were three fictitious allergic reactions (xianethis, daryosis, and chloristine). The food cues, allergic reactions, and no reaction were presented in blue, red, and green fonts, respectively. All stimuli were presented in Arial font, size 36, and were surrounded by a rectangular box drawn in the same colour as the text.
Design
All 15 foods were randomly assigned to the roles of cues: A1–A3, B1–B3, filler cues C1–C3, and filler compound cues D1E1–D3E3 (the compounds were included to increase the complexity of the task). The three fictitious allergic reactions were randomly allocated to the roles O1–O3. These roles were randomly assigned by the program at the start of the experiment. There were two training phases (see Table 1). In Phase 1, Cues A1–A3, B1–B3, and D1E1–D3E3 were presented with O1–O3, respectively, while cues C1–C3 were presented with no reaction. In Phase 2, Cues A1–A3 were presented with no reaction, while cues B1–B3 were not presented. All other contingencies were the same as those in Phase 1. All trial types in both phases were presented six times.
Design for Experiment 1
Note: Letters A1–E3 refer to food cues. Outcomes (O1–O3) refer to the fictitious allergies, xianethis, daryosis, and chloristine. Each trial type was presented six times.
Procedure
Participants were tested in a sound-attenuated cubicle. They were asked to assume the role of allergist and to attempt to determine which foods caused which allergic reactions in the hypothetical patient Mr X.
Each training trial was presented on a separate screen. The words “Mr X ate” and the relevant food(s) were shown at the top of the screen. After a 1-s delay, the words “Did Mr X experience:” and four response buttons appeared below the food(s), labelled chloristine, daryosis, xianethis, and no reaction. Participants made their predictions by using the computer mouse to click on one of the four buttons. Once a response was made, the correct outcome was highlighted (a red box appeared beneath that outcome). In addition, if participants had predicted that outcome, a tick and the word “Correct” appeared at the bottom of the screen. If a different prediction had been made, the word “Incorrect” appeared at the bottom of the screen. After a 1-s delay, a “Continue” button appeared, allowing progression to the next trial. Trials were presented in an intermixed, random order, and there was no break or discernable change between Phases 1 and 2.
On each test trial, the words “When Mr X ate” and the name of the test food were shown. After a 1-s delay, the words “What did Mr X experience?” were presented, and three response buttons (daryosis, xianethis, and chloristine) appeared below the food (no reaction was excluded as a response option at test). Once the participant chose an outcome, a scale ranging from 0 to 100 appeared beneath the three buttons, along with instructions to indicate “to what extent that food CAUSED the allergic reaction”. Once a response had been made, a “Continue” button appeared, allowing movement to the next test screen. The 15 test items were administered in a random order. Once participants had made their responses, they were not able to change their answers.
Results and discussion
For ease of explanation, the three individual exemplars of each cue class are referred to as a single cue throughout this analysis. For example, Cues A1, A2, and A3 are referred to as Cue A. A set of planned contrasts using a multivariate, repeated measures model (O'Brien & Kaiser, 1985), as well as Pearson's correlation coefficients were used to analyse the data from this and the subsequent experiment.
There was no significant effect of replication, nor did replication interact with any other effects (F < 1 in all cases). Thus, the data from the two replications were combined for the remaining analyses.
Training
At the end of Phase 1 participants generally predicted a reaction to occur following Cues A, B, and DE and no reaction to follow Cue C. At the end of Phase 1, the correct outcome was predicted on 81% of A trials, 82% of B trials, and 82% of DE + trials. No outcome was correctly predicted on 94% of C trials. Participants very quickly learned to predict “no reaction” when Cue A was presented in Phase 2; only one participant predicted that an outcome would follow the last A presentation in Phase 2.
Test
An overall accuracy score on the memory test was calculated for Cues A and B (each participant scored 0%, 33%, 66%, or 100% for each cue) and is presented in the top panel of Figure 1. Causal ratings of the A and B cues (including cues for which an incorrect choice was made on the memory measure) were also averaged for each participant and are presented in the lower panel of Figure 1. Both the memory scores and the causal ratings of the extinguished (A) cues would seem to be lower than those of the control (B) cues. Contrast analyses confirmed that memory for the outcome paired with A was lower than that for B, F (1, 59) = 6.43, p < .02, and that the causal ratings of A were lower than those of B, F(1, 59) = 156.16, p < .0001. An interaction contrast also showed that the difference between A and B on the causal rating measure was greater than that on the memory measure, F(1, 59) = 61.83, p < .0001. In an analysis of the correlation between extinction on the two measures, Cue A scores were subtracted from Cue B scores for both the percentage correct recall and mean causal judgement for each participant. This produced a measure of extinction on each of the two tests. A significant, small to moderate correlation (r = .35) was found between these difference scores.

Experiment 1: The upper panel shows the percentage of A and B cues for which the correct outcome was recalled on test. The lower panel shows the mean causal rating given to Cues A and B on test. Error bars represent standard errors of measurement.
These data are consistent with the idea that extinction results from a failure to strongly activate the outcome representation on test, as demonstrated by the decreased outcome recall and causal judgements for Cue A. The correlation between the extinction seen on the two measures further supports this interpretation. Before we accept this conclusion, however, two things should be noted. First, there is a significant interaction between cue and test type and only a small to moderate correlation between the extinction scores on the two measures. This suggests that there is some contribution to the extinction seen in causal judgements over and above that of the failure to retrieve the memories encoded in Phase 1.
Secondly, the memory task might have been influenced by the requirement to make a judgement on the causal status of the stimulus in question. That is, participants may have deliberately chosen an incorrect allergic reaction outcome in order to demonstrate their knowledge that Cue A was not the cause of the allergic reaction outcome with which it was paired in Phase 1. One way to rule out such an interpretation would be to conduct a between-subjects version of Experiment 1. Thus, participants given the memory measure would not be asked to give a causal rating of the cue. However, this would not allow an analysis of the correlation between the responses on the two measures. Experiment 2 used a different strategy to rule out any influence the causal rating might have had on the memory measure.
EXPERIMENT 2
The present experiment is different from Experiment 1 only in the nature of the test instructions. On test, participants were told explicitly that some foods were followed by no reaction in the latter stages of the experiment, but that all foods presented on test were followed by an allergic reaction in the earlier stages. They were instructed to identify this allergic reaction outcome. This more explicit instruction to indicate the earlier outcome of all cues should prevent participants from assuming that, because Cue A was not paired with an allergic reaction in Phase 2, it would be incorrect to choose the reaction paired with Cue A in Phase 1 in the memory task.
Method
The method was the same as that in Experiment 1, except in the following respects.
Participants
Participants were 57 first-year psychology students at the University of New South Wales who completed the experiment in partial fulfilment of a course requirement.
Design and stimuli
On each test trial, an additional instruction panel was presented on the left-hand side of the screen. These additional instructions read: “Please Note: This food MAY have been followed by ‘no reaction’ during the later days of Mr X's allergy tests. However, it was definitely followed by an allergic reaction during the earlier stages of the allergy tests. When this food was followed by an allergic reaction, which allergic reaction was it followed by?” The nature of these instructions necessitated that cues C1–C3 were not presented during the test trials of Experiment 2, because they had not been paired with an outcome at any stage of the experiment. Thus, Experiment 2 contained only 12 test trials.
Results and discussion
All scores were calculated and analysed as in Experiment 1.
Training
Training proceeded as expected and almost exactly as in Experiment 1.
Test
As can be seen in Figure 2, both outcome recall and causal rating were lower for the extinguished cue, A, than for the control cue, B, F(1, 56) = 5.264, p < .03, and F(1, 56) = 62.416, p < .0001, respectively. A significant interaction between cue and test type was also detected, F(1, 56) = 15.17, p < .001, showing that the difference between Cues A and B was larger on the causal judgement measure than in the memory task. As in Experiment 1, difference scores between Cues A and B were also calculated to produce extinction measures for both outcome recall and causal judgement. Again, a significant, small to moderate correlation was found between these two difference scores (r = .38).

Experiment 2: The upper panel shows the percentage of A and B cues for which the correct outcome was recalled on test. The lower panel shows the mean causal rating given to Cues A and B on test. Error bars represent standard errors of measurement.
This pattern of findings closely replicates the results of Experiment 1, confirming the finding that both causal ratings and outcome recall accuracy decline as a consequence of extinction training. Participants were explicitly instructed to disregard the cue–no-outcome trials and simply to attempt to remember the initial cue–outcome trials. Thus, the reduction in accuracy in outcome recall for extinguished cues seems to be a true failure of memory, not a misunderstanding of the requirements of the task that led participants to deliberately choose an incorrect outcome. Furthermore, the correlation between extinction in memory and that on the causal judgement measure supports the idea that extinction in causal judgements is, at least in part, due to a failure to recall the information encoded in Phase 1 of training.
GENERAL DISCUSSION
The present results show that when extinction is observed in learning, it is likely that at least part of the reduction in responding is due to a failure to recall the information encoded in the first stage of training, the cue–outcome pairings. This finding is consistent with Bouton's (1993) and Paredes-Oley and Rosas's (1999) claim that extinction is a consequence of interference between cue–outcome and cue–no-outcome memories at the point of retrieval. Another explanation is that the cue–outcome relationship learned in Phase 1 is unlearned in Phase 2 (e.g., Rescorla & Wagner, 1972). Further experiments would be required to distinguish between a retrieval interference and an unlearning account of extinction in memory.
A larger extinction effect was observed on the causal rating than on the memory measure. That is, participants gave many extinguished cues low causal ratings, despite having very good memory for the outcome with which the cues were paired in Phase 1. This finding implies an inferential contribution to the extinction seen in causal ratings (see Lovibond, 2004). Thus, for many cues, although participants recalled the cue–outcome pairings from Phase 1, they reasoned that these cues were less causal because they had not been followed by the outcome in Phase 2. We should, however, treat this evidence for an inferential account of extinction with some caution. The causal rating and memory tests no doubt differ greatly in their sensitivity, and this alone may be responsible for the differing size of the extinction effect observed on the two measures.
Bouton (1993) has argued that other phenomena, such as renewal following extinction, are also due to memory processes. In renewal, strong responding is observed if extinction takes place in a different context from that of training and test (Bouton & Bolles, 1979). This phenomenon has also been shown in human memory using a paired-associate learning task (Bilodeau & Schlosberg, 1951). It is certainly possible, therefore, that when renewal is observed in causal ratings (e.g., Vila & Rosas, 2001), it is due to a reduction in the retrieval interference normally caused by extinction training. On the inferential account, by contrast, participants reason that the cue causes the outcome except when it is presented in the extinction context. Thus, despite equivalent retrieval of the original training memories, the extinguished cue will be given a low causal rating when presented in the extinction context and a high rating when presented in the training context. The technique used in the present experiments offers a new way to test whether phenomena such as renewal result from changes in memory retrieval or from inferential reasoning processes.
