Abstract
While the concept of credibility seems like an intuitive one, research has indicated that there is no consistent definition of this construct and that credibility may, in fact, be multidimensional. This article is the first to review how the measurement of credibility in child sexual assault cases has been conducted, with the view to improve how credibility is psychometrically measured. Our findings indicate that the majority of experiments have been conducted in the United States (67%), have been based primarily on undergraduate students as participants (67%), and primarily investigated cases involving a male defendant and female victim (69%). Ultimately, among experiments investigating victim credibility, approximately 60% of all measures were based on a single item and 53% used materials not based on the testimony of the child. Moreover, credibility has been measured using a great variety of constructs such as believability, honesty, truthfulness, suggestibility, accuracy, and reliability. A more nuanced and consistent definition of credibility will be needed to facilitate meaningful applications of the research literature.
While the concept of credibility seems like an intuitive one, research has indicated that there is no consistent definition of this construct and that credibility may, in fact, be multidimensional. This article is the first to review how the measurement of credibility in child sexual assault cases has been conducted, with the view to improve how credibility is psychometrically measured. Our findings indicate that the majority of experiments have been conducted in the United States (67%), been based primarily on undergraduate students as participants (67%), and primarily investigated cases involving a male defendant and female victim (69%). Ultimately, among experiments investigating victim credibility, approximately 60% of all measures were based on a single item and 53% used materials not based on the testimony of the child. Moreover, credibility has been measured using a great variety of constructs such as believability, honesty, truthfulness, suggestibility, accuracy, and reliability. A more nuanced and consistent definition of credibility will be needed to facilitate meaningful applications of the research literature.
Witness credibility is a critical factor in jurors’ fact-finding determinations (Whobrey, Sales, & Elwork, 1981). Specifically, criminal trials often depend on witness testimony for the conveyance of disputed evidence, making assessment of witness credibility one of the most central and critical aspects of jurors’ decision-making (Kassin, 1983; Whobrey et al., 1981). The issue of perceived victim credibility is particularly relevant when it comes to cases of child sexual assault, which is said to be one of the hardest offense categories to prosecute (Wundersitz, 2003). As physical or medical evidence is often lacking in these cases, the perceived credibility of the victim becomes an essential element in the judicial process (Bottoms, Golding, Stevenson, & Yozwiak, 2007; Goodman-Delahunty, Cossins, & O’Brien, 2010). Moreover, research has indicated that even if medical evidence is present, it does not necessarily directly result in a conviction of the alleged offender (Lewis, Klettke, & Day, 2014).
However, victim credibility has been identified as having a direct influence on the outcomes of child sexual assault cases. That is, the less credible a victim is perceived to be, the less guilt is attributed to an alleged offender (Goodman-Delahunty et al., 2010; Kaufmann, Drevland, Wessel, Overskeid, & Magnussen, 2003). Yet, to date, comparatively little is known about the construct of victim credibility and how it functions in court cases.
Given the apparent centrality of jurors’ assessments of credibility to their fact-finding goal, the following question is raised: what makes a child witness credible in the eyes of the jury? In addressing this question, it is instructive to first consider definitions of credibility. For instance, according to Black’s Law Dictionary, credibility refers to the “quality in a witness which renders [his] evidence worthy of belief” (Black, n.d.). It may also be defined as “the extent to which a judge or jury believe that the witness is providing honest and accurate testimony” (Nurcombe, 1986, p. 473). Alternatively, credibility has been defined as the “ability to observe or remember facts and events about which the witness has given, is giving, or is to give evidence” ( Evidence Act 2008 [Vic], Dictionary Pt. 1). Thus, while the concept of credibility seems like an intuitive one, there is a lack of consistent definition and understanding of the concept. Indeed, understanding the factors that might contribute to the perceived credibility of child victims in sexual assault cases (e.g., victim age, gender, appearance, and behavior) is complicated by variations in the definition and assessment of credibility. That is, credibility is not a tangible or directly observable variable, but a construct that jurors perceive or infer based on the testimony provided—making the issue of measurement critical.
Early work on persuasion and attitude change provides both theoretical and operational definitions for the construct of credibility. In clarifying the meaning of the construct, Hovland, Janis, and Kelley (1953) distinguished between two components of credibility: expertness and trustworthiness. While expertness is the “extent to which a communicator is perceived to be a source of valid assertions,” trustworthiness is described as “the degree of confidence in the communicator’s intent to communicate the assertions he considers most valid” (Hovland, Janis, & Kelley, 1953, p. 21). A communicator’s credibility is said to be the weight given to his or her assertions as a function of both of these components. More recent research also suggests that credibility may be multidimensional and includes the additional subconstructs of accuracy, believability, honesty, reliability, truthfulness, suggestibility, confidence, and consistency (Brigham, 1998; Johnson & Shelley, 2014; Myers, Redlich, Goodman, Prizmich, & Imwinkelried, 1999; Pozzulo, Dempsey, Maeder, & Allen, 2010). It is unclear whether these dimensions are consistently included across the literature examining child witness credibility, though it seems that researchers have taken varied theoretical and practical approaches to measuring credibility.
The aim of this article is to review the psychometric properties underlying the measurement of perceived victim credibility in child sexual assault cases. This article has implications in terms of assisting researchers and practitioners with the conceptualization of victim credibility and enhancing the methodological rigor with which credibility is studied in future. While there is a plethora of research in this area, there is yet to be a systematic review of this literature and the measurement of credibility has been largely inconsistent. The results of this review will be presented and discussed according to the major subconstructs of credibility that emerge. The limitations of the available research will be highlighted, and the review will conclude with a discussion of implications and directions for future research.
Method
The systematic search was conducted in line with the preferred reporting items for systematic reviews and meta-analyses guidelines (Moher, Liberati, Tetzlaff, & Altman, 2009). A comprehensive search was conducted through multiple databases including PsychInfo, Medline Complete, Informit, Criminal Justice Abstracts With Full Text, Social Work Abstracts, SocIndex With Full Text, and Academic Search Complete. Relevant paper citations were hand searched for additional relevant publications. Search terms used in the review included (child* OR minor OR youth OR juvenile OR adolescent OR teen*) AND (“sexual assault” OR “sexual abuse” OR CSA OR “sex* offense” OR molest*) AND (credibility).
The following inclusion criteria were used to select articles: (1) full text publications that were available in the English language, (2) articles were peer reviewed, (3) sufficient information was presented within the methodology such that results could be extracted, (4) the focus was on child sexual assault as defined above and did not involve an adult disclosing an event from childhood, (5) participants in the research were not solely professionals such as clinicians, police officers, or social workers (given the importance of understanding layperson’s perceptions as a representation of jury eligible community members), and (6) the article reported results relevant to the perceived credibility of the victim.
The initial search returned 798 articles, with an additional 11 articles located through hand searching. Of all, 440 articles remained after the removal of duplicates. After abstract review, 62 articles were assessed as potentially meeting criteria. Following evaluation of the full text of remaining articles, 51 articles were retained for review. A summary of the systematic search process is illustrated in Figure 1.

Preferred reporting items for systematic reviews and meta-analyses flowchart.
Coding
In order to review and identify the various measures of credibility in child sexual assault cases, Table 1 reports the various constructs and scales that have been utilized. The results are reported according to the identified subconstructs of credibility, with the view to demonstrate the breadth with which credibility has been previously measured and conceptualized. Experiments were coded according to the following three rules: (1) where possible using subconstruct or dimension as identified by the researchers, (2) where constructs were not termed by the researchers, the dimension was coded according to the phrasing of the item, for example, “likelihood that the victim is telling the truth,” coded as “truthfulness” (Brigham, 1998), and (3) in cases where the authors’ terminology of a construct was deemed to be inconsistent with the phrasing of the item, the construct was coded in accordance with the phrasing of the item, for example, “how truthful the disclosure was,” coded as truthfulness, despite being termed “believability” by the researchers (Bornstein, Kaplan, & Perry, 2007). This final category occurred 10 times within the review, and these cases are identified with an asterisk within Table 1.
Measurement of Credibility.
aConstruct coded inconsistent with author’s terminology.
Coding was conducted by two of the authors. Cohen’s κ was calculated in order to determine consistency in agreement between raters regarding the terminology of constructs. The computed inter-rater reliability was Cohen’s κ = .90, which is considered to be “almost perfect” agreement (Landis & Koch, 1977).
Results
Fifty-one publications were retained for the review, reporting the results of 64 unique experiments measuring the perceived credibility of child sexual assault victims. For specificity purposes, the following results will be based on the individual experiments rather than articles.
Of all 64 experiments, 43 (67%) were conducted in the United States, 8 (12%) were conducted in Canada, and a further 7 were conducted in the UK, with 5 experiments conducted in Australia and 1 in France. Forty-three experiments (67%) used exclusively undergraduate students as participants. Forty-four experiments (69%) were based on heterosexual victim/defendant relationships, all of which involved a female victim and male alleged defendant. Thirty-four experiments (53%) used a written vignette as stimulus material on which participants based their assessment of the victim’s credibility, while a further 15 (23%) used trial transcripts. All other experiments were based on video material (11/64), real trials (2/64), or forensic interviews (2/64). Ultimately, 30 experiments (47%) provided participants with testimony given by a victim, which means that 53% of the research measuring child credibility in child sexual assault cases did not provide testimony given by the child.
With regard to the psychometric properties of the scales used, of all experiments retained in the review, 60% (78/131) of all measures, across the various subconstructs, were scored based on a single-item measure. That is, when asking respondents to rate the accuracy, believability, or honesty and so on, of the child, participants were asked to rate only one item such as the “belief of the witness [victim] testifying” (Golding, Alexander, & Stewart, 1999).
Believability
Twenty-two experiments used the subconstruct of victim believability to measure victim credibility. Only one of these publications provided a clear definition of believability as “a victim’s or defendant’s willingness to lie about the events” (Pozzulo et al., 2010, p. 53). Of the 22 experiments measuring believability, 18 (82%) used a single-item measure, such as the “believability of the victim’s testimony” (Allen & Nightingale, 1997), or the “belief of the witness [victim] testifying” (Golding et al., 1999). Alternatively, in three experiments (14%) victim believability was measured as the sum of seven items deemed to relate to believability, such as “children do not lie about sexual abuse,” “a child would probably falsely report sexual abuse to ‘go along with’ a police person or therapist who believed that the child was molested,” and “children are not capable of inventing stories of sexual abuse.” The final experiment utilized factor analysis to extract items loading onto a factor, which was subsequently interpreted as victim believability (Tubb, Wood, & Hosch, 1999). The wording of items included in this factor was not reported, however, items related to constructs such as how knowledgeable, intelligent, accurate, confident, and honest the child appeared.
Honesty
Seventeen experiments used the subconstruct of victim honesty to measure victim credibility. None of the articles retained for the review provided a definition of honesty. Of the 17 experiments, 6 (35%) used a single-item measure, for example, asking participants to rate the “honesty of the victim” (Golding, Lynch, Wasarhaley, & Keller, 2015), “the extent to which the child fabricated the allegations of child abuse” (Ross, Lindsay, & Marsil, 1999, Experiment 1), or the extent to which they agree that “the child witness is honest” (Regan & Baker, 1998).
Four experiments (24%) utilized factor analysis to extract items that loaded onto a single factor deemed to represent victim honesty. For example, Brigham (1998) extracted three items loading onto the “honesty” factor including “a child X years old is, on average [very honest/dishonest],” “in general, when compared to an adult, a X year old child is [much less/much more] likely to tell a lie,” and “how likely is a X year old child to lie about a significant event [very likely/very unlikely]?” In contrast, Ross, Jurden, Lindsay, and Keeney (2003) extracted 9 items such as “at the time [victim name] claimed her father abused her, do you think that she knew what her breasts were?,” “in general, how suggestible was [victim name]?,” and “to what extent, if any, do you believe that [victim name]’s testimony was the truth?”
A further six experiments (35%) utilized multiple-item measures, which were either summed or averaged to form a single rating for victim honesty. For example, in the three experiments by Connolly, Gagnon, and Lavoie (2008), victim honesty was measured as the average of 3 items including “how honest do you think [victim name] was?” “how sincere was [victim name]?” and “do you think [victim name] honestly believed that a sexual assault had been committed?” The remaining three experiments included items assessing honesty, sincerity, truthfulness, and the likelihood of fabrication, however, the exact wording of these items was not reported. Conversely, one further experiment utilized a Q-sort technique, asking participants to sort items relating to honesty, as being characteristic, neutral, or uncharacteristic of the child (Nunez, Kehn, & Wright, 2011).
Truthfulness
Thirteen experiments used the subconstruct of victim truthfulness to measure victim credibility, with only one providing a definition by stating that “truthfulness refers to a victim’s or defendant honesty” (Pozzulo et al., 2010, p. 53). Based on this definition, it is not clear whether truthfulness is in fact a separate subconstruct to that of honesty. Of the 13 experiments measuring truthfulness, 8 experiments (62%) used a single-item measure asking participants to rate for example “how truthful the disclosure was” (Bornstein et al., 2007), the “likelihood that [the victim] is telling the truth” (Brigham, 1998), or “how truthful do you find the alleged victim’s testimony?” (Pozzulo et al., 2010). A further two experiments (15%) used 2 separate items, one relating to the victim telling the truth and other one relating to the defendant telling the truth, with these items subsequently combined (O’Donohue & O’Hare, 1997; O’Donohue, Smith, & Schewe, 1998). Two other experiments (15%) used factor analysis to reveal items loading onto the factor “victim truthfulness,” with one experiment extracting 7 items (Schmidt & Brigham, 1996, Experiment 1) and the other one extracting 5 items (Schmidt & Brigham, 1996, Experiment 2). In both of these experiments, the wording of items included was not reported. The remaining experiment used the average of 5 items relating to whether the victim was honest, truthful, believable, trustworthy, and convincing, however, again the wording of these items was not reported.
Suggestibility
Thirteen experiments used the subconstruct of victim suggestibility to measure victim credibility, which has been defined as “whether the child’s story was suggested to him by other persons” (Bottoms & Goodman, 1994, p. 720). Eleven experiments (85%) utilized a single-item measure for example asking participants to rate the “suggestibility of the child witness” (Ross et al., 1999), however, the majority of these experiments (7/11) did not report the wording of this item. The remaining two experiments (15%) measured suggestibility using factor analysis. Bottoms and Goodman (1994) extracted two items relating to suggestibility from the child’s mother and from the police officer, however, the wording of these items was not reported. Similarly, Brigham (1998) extracted three items loading onto the factor referred to as suggestibility. These items included “how suggestible do you believe a(n) X year-old child is, compared to adults? [much more suggestible than an adult/much less suggestible than an adult]” “when threatened with harm to self or family members, a(n) X year-old child would be [much more/much less] likely than an adult would be to change his/her story,” and “can a child X years old be coached by adults to lie successfully, so that other adults will be fooled? [certainly/definitely not].”
Accuracy
Twelve experiments used the subconstruct of victim accuracy to measure victim credibility, with one experiment providing a definition of accuracy as being “the degree to which the statements were consistent with what actually occurred” (Pozzulo et al., 2010, p. 53). Six experiments (50%) utilized a single-item measure of victim accuracy asking participants to rate, for example, the “accuracy of the child’s memory for the specific acts claimed to be sexual abuse” (Ross et al., 1999) “how accurate do you find the alleged victim’s testimony?” (Pozzulo et al., 2010), or “in your opinion how likely is it that while testifying in court, the main child was accurate about being sexually abused?” (Myers et al., 1999).
A further two experiments (17%) utilized factor analysis to extract items loading onto “victim accuracy.” Schmidt and Brigham (1996) extracted 3 items, however, the wording of these items was not reported. Brigham (1998) also extracted 3 items including “compared to adults, a child X years old is, on average, able to recall traumatic or stressful events [much more accurately/much less accurately],” “when describing to an adult an event that happened to him/herself, an X year-old child is likely to be [much more concerned/much less concerned] that the description is completely accurate than an adult would be,” and “compared to an adult, a child X years old is, on average, when recalling an important event that occurred 6 months age [much more accurate/much less accurate].”
The final four experiments (33%) scored the average of multiple items as the measure of victim accuracy. McAuliff, Lapin, and Michel (2015) averaged the scores of 7 items relating to accuracy, reliability, credibility, clarity, consistency, certainty, and how well-spoken the victim was, however, the wording of these items was not provided. The three experiments by Connolly et al. (2008), scored victim accuracy as the average of 4 items including “how intelligent do you think [victim name] was?” “how accurately do you think [victim name] recalled the details of the event?” “how well did [victim name] understand the events she described?” and “[victim name] was asked seven particular questions about the event. How many of those questions do you think she answered accurately?”
Reliability
Eight experiments measured victim credibility via the subconstruct of victim reliability. Reliability was defined as “the degree to which a juror can depend on the statements made by the victim or defendant” (Pozzulo et al., 2010, p. 53). Six experiments (75%) utilized a single-item measure of reliability asking participants to rate, for example, the degree to which they agree that the “child witness is reliable” (Regan & Baker, 1998) or “how reliable do you find the alleged victim’s testimony?” (Pozzulo et al., 2010). The remaining two experiments analyzed the data using factor analysis, extracting 9 items relating to victim reliability including such things as “how likely is it that the child was lying?” “how suggestible did the child seem?” and “how likely is it that the child misunderstood the defendants actions?” (Castelli, Goodman, & Ghetti, 2005, Experiments 1 and 2).
Consistency
Five experiments utilized the subconstruct of victim consistency, however, no definition of consistency was provided in any of the retained experiments. In all five experiments, victim consistency was measured utilizing a single item. The wording of this item was not reported in four experiments, while the final experiment asked participants “how consistent was the child’s testimony?” (Cashmore & Trimboli, 2006).
False Belief
Four experiments used the subconstruct of whether the child held a false belief as the measure of victim credibility, with no definition provided in any of the experiments. In all four of these experiments, a single-item measure was utilized asking participants to rate, for example, “whether the victim honestly believed the abuse charge” (Bottoms, Nysse-Carris, Harris, & Tyda, 2003) or how strongly they agreed that the “child witness has falsely remembered the encounter” (Regan & Baker, 1998).
Confidence
Three experiments measured victim credibility via the subconstruct of victim confidence, with no definition provided in any of the retained experiments. In all three of these experiments, victim confidence was measured using a single-item measure, with the wording of this item only reported in one experiment. Specifically, Myers, Redlich, Goodman, Prizmich, and Imwinkelried (1999) measured victim confidence by asking participants “how confident did the child seem to be while the child testified?”
Full Disclosure
A further two experiments assessed victim credibility as whether the victim provided a full disclosure of the events. In these two experiments by Castelli, Goodman, and Ghetti (2005), factor analysis was used, extracting 4 items relating to whether a full disclosure was provided. These items included the following questions: “how likely is it that the child forgot to tell things that really happened?” “how likely is it that the child told part of what happened about abuse, but not all of what happened?” “how likely is it that the child refused to tell things that really happened?” and “how likely is it that the child was too frightened to tell things that did happen?”
Convincing
One experiment measured victim credibility via the subconstruct of whether the victim was viewed as being convincing. Cashmore and Trimboli (2006) utilized a single-item measure, asking participants “how convincing was the child’s testimony?”
Credibility
Finally, 31 experiments measured victim credibility directly. Credibility has been broadly defined as “whether the child was accurate and truthful in [his] actual testimony” (Bottoms & Goodman, 1994), “the extent to which a judge or jury believe that [she] is providing honest and accurate testimony” (Regan & Baker, 1998), “the jurors perception of the victim’s or defendant’s likelihood of telling the truth as he or she knows it” (Pozzulo et al., 2010), or “the perceived memory performance and honesty of the victim” (Ross et al., 2003).
Of the 31 experiments measuring victim credibility, 10 experiments (32%) utilized a single-item measure asking participants, for example, “do you think that [victim name’s] testimony is credible?” (Esnard & Dumas, 2013), “how credible was the victim” (Klettke, Graesser, & Powell, 2010), or “how credible do you find the alleged victim’s testimony?” (Pozzulo et al., 2010). A further nine experiments (29%) used factor analysis to extract items loading onto the factor “victim credibility.” For example, Rogers, Titterington, and Davies (2009) extracted 5 items, including “[victim name] was accurate at giving evidence of the incident,” “[victim name] remembers the events of the incident clearly,” “[victim name] is telling the truth about what happened,” “[victim name] is a dependable witness,” and “[victim name]’s statement to the police is credible.”
The remaining 12 experiments (39%) used multiple items, which were either summed or averaged to obtain a single credibility score. The majority of these experiments (9/12) did not provide a full list of the items included in this scale. In the three experiments by Connolly et al. (2008), victim credibility was scored as the average of two items including “how much weight did you give to [victim name’s] testimony? In other words, to what extent did [victim name’s] testimony influence your decision about what happened?” and “overall, how credible was [victim name]?.”
Discussion and Implications
This systematic review investigated how previous research has measured the variable “perceived victim credibility” in cases of child sexual assault. The review highlights that experiments vary greatly in their measurement of credibility and that there are considerable methodological limitations to the research as a whole. There were several findings.
Firstly, and most importantly, there is no standardized measure of credibility utilized in any of the reviewed studies. Specifically, measures of credibility such as believability, truthfulness, and credibility have been used interchangeably, limiting the generalizability of findings across studies. While the terminology varies greatly, definitions are rarely provided to participants, who are left to create their own interpretation when responding to attitudinal items. However, whether credibility is a complex construct encompassing multiple subconstructs such as believability and truthfulness or can be accurately and validly measured and conceptualized as a single construct is currently an empirical question that needs to be addressed.
Secondly, the majority of studies utilized a single-item measure of credibility. Thus, a single-item measure asking participants to rate, for example, the truthfulness of the victim’s testimony is measuring only one subconstruct of credibility, and therefore conclusions should not be extrapolated to the whole construct of victim credibility. Even a single-item measure asking participants to rate the victim’s “credibility” is methodologically limited. The problem with a single-item measure of a construct is that it does not allow for the measurement of internal consistency, containing considerable measurement error (Loo, 2002). Single-item measures also lack precision and scope in their measurement of a construct and do not allow for the broad spectrum of factors that might contribute to such a multidimensional construct (Gliem & Gliem, 2003). While single-item measures are of some use and should not be rejected in their entirety, no single item is sufficient to provide a complete measure of such a complex concept as credibility. In order to minimize the effects of error variance and improve reliability of the measurement of credibility, a multi-item scale should be utilized. It is conceivable that the lack of consistent methodology and measurement of credibility has resulted in inconsistencies in the results of available research to date.
Thirdly, perceived credibility has largely been measured after participants have been asked to read a vignette depicting an alleged offense rather than by providing examples of actual testimony provided by the child. It is likely that the credibility of the child witness is largely based on the substantive evidence provided by that witness. Therefore, asking participants to rate the credibility of a witness without providing examples of such evidence is almost an impossible task and sheds light on how future research pertaining to child witness credibility can and should be improved.
Fourthly, the majority of studies to date have investigated only heterosexual victim/defendant relationships, all of which have explored cases involving a female victim and male defendant. While the majority of victims of child sexual assault are female (approximately 70% female and 30% male self-report victimization), it is possible that this is reflective of underreporting of male victimization (Stoltenborgh, van IJzendoorn, Euser, & Bakermans-Kranenburg, 2011; Wundersitz, 2003). It is critical that research exploring child sexual assault continues to explore cases involving male victims and female defendants as well as same sex cases, avoiding any possible contribution to increasing myths regarding the occurrence of such offenses and improving our understanding of victim credibility in such cases.
Another observation is that the large majority of studies solely utilized undergraduate students as participants. As compared to the general public, undergraduate students have been shown to have less solidly formed attitudes, a weaker sense of self and stronger cognitive reasoning abilities (Sears, 1986). Thus, findings based on such a narrow sample should be interpreted with some caution. In addition, approximately 80% of studies were conducted within either the United States or Canada. Although this is due to the current review being limited to papers written in the English language, only a small number of studies have been conducted in other English speaking countries (e.g., the UK and Australia). There is considerable scope to extend research in both Western and non-Western countries.
The present study systematically reviewed the measurement of perceived victim credibility, highlighting the lack of methodological rigor and consistency in the available research to date. Ultimately, in order to improve procedural fairness and safeguard the rights of both the accused and the victim, it is important to understand the factors influencing perceived victim credibility. Most notably, to improve the methodological rigor in this area, there is a strong need for a valid and reliable scale for the measurement of perceived victim credibility. It is further recommended that research be extended beyond the United States and Canada, that further studies do not utilize exclusively undergraduate students as participants, and that researchers explore male/female and same sex victim/defendant relationships. However, before research investigating the impact of extralegal factors can be extended, a clear definition and consistent measurement of credibility is required.
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
