Abstract
Background:
The Delirium Observation Screening Scale (DOS) is designed to detect delirium by nurses’ observations and has shown good psychometric properties. Its use in palliative care unit patients has not been studied.
Aim:
To determine diagnostic and concurrent validity, internal consistency, and user-friendliness of the Delirium Observation Screening Scale administered by bedside nurses in palliative care unit patients.
Design:
In this descriptive study, psychometric properties of the Delirium Observation Screening Scale were tested by comparing the performance on the Delirium Observation Screening Scale (bedside nurses) to the algorithm of the Confusion Assessment Method and the Delirium Index (DI) (researchers). Paired observations were collected on three time points. Afterward, the user-friendliness of the Delirium Observation Screening Scale was determined by bedside nurses using a questionnaire.
Setting/participants:
In total, 48 patients were recruited from one palliative care unit (PCU) of a university hospital. Of the 14 eligible bedside nurses of the palliative care unit, 10 participated in the study.
Results:
Delirium was present in 22.9% of patients. Diagnostic validity of the Delirium Observation Screening Scale was very good (area under the curve = 0.933), with 81.8% sensitivity, 96.1% specificity, 69.2% positive, and 98% negative predictive value. Concurrent validity of the Delirium Observation Screening Scale with the Delirium Index was moderate (rSpearman = 0.53, p = 0.001). The Cronbach’s alpha for all Delirium Observation Screening Scale shift scores was 0.772. Generally, bedside nurses experienced the Delirium Observation Screening Scale as user-friendly. However, most Delirium Observation Screening Scale items (n = 11/13 items) need verbally active patients to perform the observations correctly.
Conclusion:
The Delirium Observation Screening Scale can be used for delirium screening in verbally active palliative care unit patients. The scale was rated as easy to use and relevant. Further validation studies in this population are required.
Keywords
Background
Delirium is a common disorder in palliative care inpatients, characterized by disturbance of consciousness, change in cognition, or development of a perceptual disturbance that occurs over a short period of time and tends to fluctuate over the course of the day.1–4 Recognition and appropriate management of delirium in palliative care are crucial because the syndrome has negative effects on patients’ and proxies’ quality of life and interferes with the provided care.5–8 Unfortunately, delirium remains often unrecognized by clinicians and is thus inadequately or undertreated.1,9,10 Therefore, the development of screening tools for improving delirium recognition has been extensively studied.11–13
A recent systematic review identified 11 bedside delirium screening scales. 14 Considering their test performance, ease of use, and brevity, the authors found best evidence to support the use of the Confusion Assessment Method (CAM). However, its performance varies depending on the skills and discipline of the examiner.14–16 When used for surveillance by bedside nurses in the real-life clinical practice, the accuracy of the CAM is poor. 15 Time required for extensive training and correct administration to achieve valid CAM assessments poses high burden and thereby limits the usefulness for bedside nursing. 17 However, nurses’ clinical observations play an important role in the early recognition and monitoring of delirium. Therefore, other tools are needed for screening, which are based on bedside observations of behavior and which can be integrated easily into daily routine care without undue response burden.18–20
One of the scales described in the mentioned review 14 that meets these criteria is the Delirium Observation Screening (DOS) Scale. 21 This tool has been tested in various hospital populations and can be regarded as reliable and valid for detection and measuring severity of delirium by nurses’ observations during routine care.21–24 Its ease of use and relevance for practice and the absence of response burden on patients make this scale eligible to implement in daily care.21–23 Yet its use in the palliative care unit (PCU) population has not been studied.
The aim of this study was to examine the diagnostic and concurrent validity and internal consistency of the DOS when applied by bedside nurses in PCU patients. In addition, its user-friendliness in monitoring this patient group was described.
Methods
Design, setting, and population
A prospective study was conducted in a PCU of a university hospital. Patients aged 18 years or older, Dutch-speaking, and verbally testable who were consecutively admitted to the PCU (November 2009–June 2010) were recruited by the PCU psychologist within 24 h of admission. Patients admitted in the imminent terminal stage of life (terminal sedated/comatose) were excluded. Written/proxy informed consent was obtained. At the end of the study, bedside PCU nurses were recruited to evaluate the user-friendliness of the DOS. Nurses who never filled out a DOS were excluded from this usability evaluation part of the study. The study was approved by the Medical Ethics Committee of the Leuven University Hospitals.
Delirium assessment
Delirium was independently evaluated during the first 10 days of the patients’ stay at the PCU by bedside nurses and one of the three researchers (M.D., N.B., and M.P.), both blinded to each others’ ratings. Bedside nurses used the DOS Scale 21 to rate delirium on a daily basis. The assessments were performed in enrolled patients three times a day at the end of each 8-h shift. The DOS contains 13 observations of behavior, each scored as absent, present, or unable. Total scores range between 0 and 13 for each 8-h shift, in which unable ratings are scored as 0. The total day score (24 h) is the mean of the three shift scores, with 13 as the highest possible day score. A score of 3 or more indicates delirium. 23
The researchers performed a maximum of three assessments in enrolled patients on three different days. These assessments were randomly chosen within the same 8-h shift (morning or evening shift) of the bedside nurses’ assessments and included completion of the diagnostic algorithm of the CAM25,26 and the Delirium Index (DI). 27 According to the CAM algorithm, the criteria acute onset, fluctuation, inattention, and disorganized thinking or altered level of consciousness have to be positive for a diagnosis of delirium. The DI is a delirium severity tool with 7 items scored on a scale ranging from 0 (absent) to 3 (present and severe). Total score ranges from 0 to 21, in which a higher score indicates greater severity. The CAM algorithm and DI were completed after a structured cognitive assessment, which included the items “orientation in time and place,” “immediate recall,” and “short-term verbal memory” of the Mini-Mental State Examination (MMSE); 28 an attention test (e.g. Attention Screening Examination); 29 and questions to nurses or relatives about the acute onset of symptoms. 26
Before the start of the study, bedside nurses and researchers were trained in performing the instruments by two research investigators (E.D. and K.M.), both having extensive research and clinical expertise in delirium. Researchers were trained according to criteria set in the manuals of CAM 26 and DI, 27 including evaluation of four clinical cases and follow-up discussions. Interrater reliability of the researchers, calculated two by two in a random sample of seven paired observations of enrolled patients, was κ = 1.00 (p < 0.001) for the CAM and DI. Bedside nurses were educated in the use of the DOS 21 during a 1-h course. The interpretation of DOS items was explained, and an instruction form was added to each DOS.
User-friendliness of the DOS
At the end of the study, nurses had to complete a 25-item “usability” questionnaire, which was adapted from Van Gemert and Schuurmans. 23 In total, 23 items are scored on a four-point Likert scale (strongly disagree/mainly disagree/mainly agree/strongly agree). The questionnaire assesses the content clarity of the scale (n = 4 questions), its relevance and feasibility for practice (n = 2 questions), and the clarity of DOS items (n =13 questions), and it evaluates nurses’ perception of their competence necessary to fill out the scale (n = 4 questions). An additional question about time to complete the DOS and an open question “Any other comments” were added.
Statistical analysis
Data were analyzed using SPSS version 17.0. Descriptive analyses were performed to summarize the characteristics of patients and nurses and the results of the user-friendliness of the DOS.
Paired delirium ratings of bedside nurses and researchers were compared to explore the diagnostic validity of the DOS for the CAM algorithm, their level of agreement, and the concurrent validity between the DOS and DI. Since CAM/DI assessments were only available for morning or evening shifts, only DOS shift scores were included in these analyses.
Diagnostic validity of the DOS was examined by constructing a receiver operating characteristic (ROC) curve and by calculating sensitivity, specificity, and positive and negative predictive values for different cutoff points of the DOS shift scores. Classification of patients as “delirious” (positive CAM and DOS shift scores ≥ 3) and “nondelirious” (negative CAM and DOS shift scores < 3) was further tested by performing agreement statistics (proportion of observed agreement (P0) and Cohen’s kappa coefficients (κ)), in combination with the prevalence and bias index. 30 Moreover, P0 is the proportion of exact agreement between two assessment methods, while κ corrects for chance. Paradoxes in the values of P0 and κ can occur because of prevalence and bias effects.30–32 First, the stability of κ is influenced by the variability of the sample (i.e. the prevalence of positive or negative ratings) and will be reduced if the ratings are homogeneous, indicated by the prevalence index. 30 Second, the κ can be influenced by a bias effect, indicated by the bias index, 30 which occurs when disagreement between the assessment methods is asymmetrical. A large bias index reflects a tendency of a systematically different disagreement between the two methods, affecting the interpretation of the κ, which will be higher than when bias is low or absent. To explore concurrent validity between DOS shift scores and total DI scores, the Spearman’s rho correlation coefficient was used. Correlations were calculated for the total group and for the delirious group. Additionally, internal consistency of the DOS was calculated based on all DOS shift scores together using the Cronbach’s alpha and item-total correlations.
Results
Sample
A total of 98 patients were admitted to the PCU, of whom 14 refused to participate, and 36 patients were excluded because participation was too burdensome according to the researchers’ opinion (n = 1), because of death or comatose state before study involvement (n = 12), or because of inability to communicate (n = 23). Admission characteristics of the 48 included patients are shown in Table 1. Patients excluded or who refused to participate did not differ significantly from those included in terms of gender (men, n = 26/50, 52% versus n = 30/48, 62.5%; p = 0.315) and age (median 76 ((interquartile range (IQR) = 17) versus 72 (IQR = 11); p = 0.248).
Admission characteristics of included patients (n=48).
Q1: first quartile; Q3: third quartile; COPD: Chronic Obstructive Pulmonary Disease.
A maximum of 1440 DOS (= 48 × 3 × 10) and 144 CAM ( = 48 × 3) observations were expected to be completed. However, because of terminal state or death of included patients during study participation, only 1108 DOS and 123 CAM observations were performed, generating 113 paired observations. For the other 10 observations, delirium measurements by bedside nurses and researchers were not made during the same 8-h shift. In these paired observations, all DOS items were rated. In 6%, 1 to 3 items were rated as “unable” to score.
Of the 17 bedside nurses, 14 were eligible for DOS usability evaluation (2 on maternity leave and 1 newly employed who never filled out a DOS); 10 of them returned the questionnaire (response rate = 71.4%). Nurses’ mean age was 44.2 years (standard deviation (SD) = 8.9 years). Their mean number of work experience as a nurse in general was 22.4 years (SD = 9.6 years) of which 9.1 years (SD = 2.2 years) with palliative care patients. Most nurses were female (n = 8/10), had bachelor’s degree (n = 7/10), and received delirium training for the last 5 years (n = 9/10).
Occurrence rates of delirium
Delirium (at least one positive CAM score) was present in 11 of the 48 patients (22.9%) or in 11 of the 113 paired observations (9.7%). An overall DOS-shift score of 3 or more occurred in 131 of the 1108 DOS observations (11.8%), indicating possible delirium.
Diagnostic and concurrent validity
The DOS showed an area under the ROC curve (AUC) of 0.933 (95% confidence interval (CI): 0.819–1.000) (Figure 1). The original cutoff point of 3 can be considered as good. Bedside nurses identified nine true-positive delirium observations and only two false-negative observations. Of the 102 observations, 4 without delirium were false positive. This results in a sensitivity of 81.8% and specificity of 96.1%. An acceptable positive predictive value and high negative predictive value were demonstrated in Table 2. Agreement between the DOS and CAM in detecting delirious and nondelirious patients was also good (P0 = 0.947; κ = 0.721, 95% CI: 0.509–0.932, p < 0.001). The bias and prevalence index were 0.02 and 0.79, respectively. Concurrent validity of paired DOS shift scores with total DI scores was moderate (rSpearman = 0.53; p < 0.001). The mean DI score for observations with a DOS shift score of 2 or lower was significantly lower than for observations with a DOS shift score of 3 or more (3.16 (SD = 2.899) versus 10.08 (SD = 3.475); p < 0.001). For the delirious group (13 paired observations), the correlation coefficient between the DOS and DI was 0.73 (p < 0.01).

ROC curve of DOS shift scores with the CAM as reference standard.
Comparison of delirium ratings between bedside nurses (DOS) and researchers (CAM) in 113 paired observations.
DOS: Delirium Observation Screening Scale; CAM: Confusion Assessment Method; CI: confidence interval.
Sensitivity = 81.8% (95% CI = 52–95); specificity = 96.1% (95% CI = 90–98); positive predictive value = 69.2% (95% CI = 42–87); negative predictive value = 98% (95% CI = 93–99); diagnostic accuracy = 94.7% (95% CI = 89–98).
Internal consistency (Table 3)
The Cronbach’s alpha coefficient for all DOS shift scores was 0.772. For item-total correlations, most items (e.g. items 1, 2, 4, 5, 6, 7, 8, 9, and 10) correlated moderately (rPearson = 0.566–0.401) and fairly (items 3, 11, and 13) (rPearson = 0.390–0.254) with the sum of the other items, while item 12 correlated weakly (rPearson = 0.177) (Table 3).
Pearson item-total correlation coefficients of the DOS (n=48 patients, 1108 test occasions).
DOS: Delirium Observation Screening Scale; IV: intravenous.
User-friendliness (Table 4)
All respondents (n = 10) mainly/entirely agreed that the concepts of the DOS items are clear, compatible with the language used in practice, and free of values and judgment. The majority (n = 9) further agreed that differences in the response options are mainly/entirely clear. Agreement about clarity (n = 9) is further reflected in all single-DOS items (except for items 2 and 6 for which one nurse mainly disagrees). All nurses mainly/entirely agreed that they had sufficient knowledge from training and experience to evaluate the observations on the scale. However, still one nurse said that she required help from others to rate the DOS, and one nurse disagreed that the instructions helped in choosing the correct answers. Most nurses mainly/entirely agreed that the DOS is a handy instrument (n = 9) and adds value to their nursing practice (n = 9). Finally, the median time to score the DOS was 1 min (IQR = 1) (Table 4).
Ease of use of the DOS (n=10 bedside nurses of the palliative care unit).
1 missing value; DOS: Delirium Observation Screening Scale; IV: intravenous.
Discussion
To our knowledge, this is the first study examining the diagnostic and concurrent validity, internal consistency, and user-friendliness of the DOS administered by bedside nurses in a PCU. The good diagnostic values of the DOS observed in surgical and geriatric populations (sensitivity = 89%–100%, specificity = 76%–96.6%)21–23 and its ease of use in surgical patients 23 could be confirmed in PCU patients.
The DOS discriminates very well between delirious and nondelirious patients, with an AUC of 0.933, as compared to the CAM as reference standard. Although the sensitivity rate (81.8%) was somewhat lower than reported in earlier studies,21–23 this result is still acceptable. More importantly, there were only two false-negative observations. The positive predictive value or the proportion of delirious patients correctly diagnosed as delirious was good and in line with the previous findings (47%–88.9%).21–23 The negative- predictive value was high, indicating that delirium was rarely present with a DOS shift score lower than threshold 3. This good diagnostic validity of the scale is confirmed by a substantial agreement between the DOS and CAM, tested with kappa statistics. However, the magnitude of the κ coefficient may be reduced because of the prevalence effect, revealing that κ was influenced by homogeneity of the sample. Yet the κ was not affected by a systematically different classification pattern between the two instruments (bias index = 0.06).
Concurrent validity of the DOS with the DI was moderate but still acceptable. Subgroup analysis with only delirious patients increased the correlation between both scales, suggesting that the DOS is valuable for monitoring delirium severity in delirious PCU patients. In the study of Scheffer et al., 24 where the DOS was compared with the Delirium Rating Scale–Revised-98, 33 a slightly stronger correlation was found (rPearson = 0.67). However, the use of a different statistical test (e.g. Pearson correlation) can clarify this discrepancy because our result was similar when this test was used (rPearson = 0.68).
Reliability analysis showed good internal consistency. Only the item-total correlation for DOS item “is easily or suddenly emotional” was low, but deleting the item did not change the internal consistency more than 0.002.
In line with Van Gemert and Schuurmans, 23 PCU nurses evaluated the user-friendliness of the DOS generally as good. Despite the small sample size (n = 10), some valuable comments on the individual DOS items were highlighted. Looking at the nurses’ ratings on clarity of these 13 items, none of them were found to be entirely clear for all nurses. Group discussions with the nurses revealed that the perceived difficulties with DOS items were not related with the used concepts themselves, but with the setting of palliative care. For example, some observations on the scale may mimic typical symptoms of advanced illness in palliative care (e.g. emotional, slower reaction, high levels of fatigue), which makes scoring sometimes difficult. Furthermore, most items require that patients are verbally active in order to make observations, indicating that it is difficult to use the DOS in patients in the imminent terminal stage of life. Therefore, nurses suggested an adaptation to improve usability of the scale; for example, to add an extra section with the specific reason why assessment is impossible. Further research is warranted to investigate these adaptations.
Despite these comments, our findings suggest that the DOS and its original threshold can be validly and reliably used for detection and monitoring of delirium severity by bedside nurses in the PCU population. Because of its time-efficiency and ease of use, the DOS can easily be implemented in daily practice, which is an important step in improving the detection of delirium. 34
This study has some limitations. First, only half of the patients (n = 48/98) admitted to the PCU were enrolled in the study. However, no significant differences in gender and age were found between the included and nonincluded patients. Moreover, this recruitment problem is in line with previous studies, where difficulties in recruiting PCU patients to research are well described.16,35 Second, the reference standard for diagnosing delirium may be criticized, because it was the CAM algorithm evaluated by researchers instead of the Diagnostic and Statistical Manual of Mental Disorders (4th ed.; DSM-IV) criteria scored by an experienced physician. Nevertheless, the reliability of the reference standard was guaranteed as the researchers were extensively trained by two experts in delirium using a validated diagnostic model that we successfully used in previous studies.15,36,37 Moreover, a recent study shows that the performance of the CAM algorithm proved well against the DSM-IV criteria in the hands of experienced clinicians. 38 Third, the validity analyses were based on 113 paired observations in 48 patients, implying that these observations were not independent, which could have potentially influenced the results. However, the main objective of this study was descriptive, not inferential, and this is not expected to be substantially affected by nonindependence. Moreover, our findings concur with previous studies on validity of the scale.21–23 Finally, paired delirium ratings by bedside nurses and researchers were not conducted at the same moment in time. This could have biased the results, given the fluctuating course of delirium throughout the day. However, measuring delirium simultaneously was not possible, because of differences in the scoring methods of the instruments used; DOS scores are based on observations made in the previous 8 h, and scoring of the CAM/DI is based on observations made at one moment in time, extended by others (e.g. involving relatives’/nurses’ observations for acute onset and fluctuation aspects). As a consequence, we tried to minimize the time span by using only evaluations performed within the same 8-hour shift in the analyses.
In conclusion, delirium detection in PCU patients suffering from symptoms of advanced illness is challenging. The DOS offers bedside nurses a promising tool for screening and monitoring delirium and its severity in this population. The scale is easy to use in verbally active PCU patients (e.g. scoring requires no extensive training) and is useful in nursing practice (e.g. to score in about 1 min). However, further validation studies in this specific population are required to confirm the results of this study.
Footnotes
Acknowledgements
The authors would like to thank all nurses and participating patients of the palliative care unit of the University Hospitals Leuven. Special thanks to Rita Van Nuffelen, head of nurses of the palliative care unit, and Nathalie Wellens, MSc SLP, PhD.
Declaration of conflicting interests
The authors declare that there is no conflict of interest.
Funding
This research received no specific grant from any funding agency in the public, commercial, or not-for-profit sectors.
