Erin ChaseORCID, Nicole Moreira, Brittany E Blanchard , [...]
View All
Abstract
Introduction:
Although people with mental health disorders are more likely to die by suicide, individuals experiencing suicidality are frequently excluded from clinical trials of mental health treatment due to safety and liability concerns. This approach limits the generalizability of trial results and opportunities for intervention. This descriptive study aimed to report outcomes and lessons learned for a suicide risk management protocol implemented for participants reporting suicidal ideation in a comparative effectiveness clinical trial that enrolled patients screening positive for posttraumatic stress disorder or bipolar disorder. Specifically, we examined the proportion of trial participants reporting suicidal ideation, their chosen risk management plan, suicide attempts, and death by suicide. Also, because few studies have examined whether the survey modality of suicide screening impacts endorsement rates, we compared suicide ideation endorsement, patient demographics, and chosen risk management plans across phone and web survey modalities.
Methods:
Descriptive statistics were used to report the proportion of participants in the comparative effectiveness trial who reported suicidal ideation and activated the suicide risk management protocol, as well as the chosen risk management plans for those with active suicidal ideation. Chi-square tests of independence and Fisher’s exact tests were used to test for differences in demographics, screening question responses, and chosen risk management plans, respectively, between web versus phone survey modalities among those that activated the suicide risk management protocol.
Results:
Of the 1004 participants in the trial, 72% endorsed current suicidal ideation or previous suicidal behavior at baseline and activated the study’s suicide risk management protocol. There were two suicide attempts in the sample (0.28%), and one of which resulted in death (0.14%). There were no statistically significant differences in SRMP activation between phone and web-based survey modalities. Among participants who activated the suicide risk management protocol and endorsed active suicidal ideation, selection of risk management plans did not vary by survey modality. Participants most frequently opted to visit their community health center (42%) or to call the National Suicide Prevention Lifeline (32%) as their chosen risk management plan.
Discussion:
We developed and implemented the suicide risk management protocol for a multisite clinical trial enrolling patients with complex mental health conditions. Although a higher proportion of participants activated the SRMP compared to previous trials, rates of suicide attempts and suicide deaths were low. Our findings indicated no differences in positive screening rates among trial participants and no differences in safety plan selection by survey modality among participants entering the SRMP. This suggests that similar protocols may be used to screen for and manage suicidality in clinical trials, and protocols can be administered via phone and web-based surveys.
Research article
Available accessResearch articleFirst published April, 2026pp. 145-154
Robert Y LeeORCID, Kevin S LiORCID, James Sibley , [...]
View All
Abstract
Background:
Natural language processing allows efficient extraction of clinical variables and outcomes from electronic health records (EHRs). However, measuring pragmatic clinical trial outcomes may demand accuracy that exceeds natural language processing performance. Combining natural language processing with human adjudication can address this gap, yet few software solutions support such workflows. We developed a modular, scalable system for natural language processing-screened human abstraction to measure the primary outcomes of two clinical trials.
Methods:
In two clinical trials of hospitalized patients with serious illness, a deep-learning natural language processing model screened electronic health record passages for documented goals-of-care discussions. Screen-positive passages were referred for human adjudication using a REDCap-based system to measure the trial outcomes. Dynamic pooling of passages using structured query language within the REDCap database reduced unnecessary abstraction while ensuring data completeness.
Results:
In the first trial (N = 2512), natural language processing identified 22,187 screen-positive passages (0.8%) from 2.6 million electronic health record passages. Human reviewers adjudicated 7494 passages over 34.3 abstractor-hours to measure the cumulative incidence and time to first documented goals-of-care discussion for all patients with 92.6% patient-level sensitivity. In the second trial (N = 617), natural language processing identified 8952 screen-positive passages (1.6%) from 559,596 passages at a threshold with near-100% sensitivity. Human reviewers adjudicated 3509 passages over 27.9 abstractor-hours to measure the same outcome for all patients.
Discussion:
We present the design and source code for a scalable and efficient pipeline for measuring complex electronic health record-derived outcomes using natural language processing-screened human abstraction. This implementation is adaptable to diverse research needs, and its modular pipeline represents a practical middle ground between custom software and commercial platforms.
Research article
Available accessResearch articleFirst published April, 2026pp. 155-165
Jessica RoydhouseORCID, Anne ZolaORCID, Monique Breslin , [...]
View All
Abstract
Background:
There is growing recognition of the importance of patient-reported tolerability in complementing traditional clinician-reported safety evaluation of cancer therapies. Recent regulatory guidance listed the evaluation of overall side effect impact as a core patient-reported outcome in oncology clinical trials. A single item (‘GP5’) that asks about side effect bother is included in the Functional Assessment of Chronic Illness Therapy and has been used to capture overall side effect impact. This paper sought to expand the evidence base for GP5 by examining its association with clinician-reported treatment-emergent adverse events and patient-reported global health.
Methods:
We examined six commercial cancer clinical trials that collected GP5. The patient population was drawn from the safety population and the analysis focused on the first on-treatment assessment. Clinician-reported adverse events were classified as symptomatic if such adverse events were considered amenable to patient self-reporting (e.g. nausea). Chi-square tests and Pearson’s correlation were used to examine associations. We considered adverse event grade and frequency, both for symptomatic adverse events and any type of adverse events. Global health was measured using the visual analogue scale of the EuroQol-5 Dimensions-3 Levels measure. ‘Moderate-severe’ bother was characterised as scores of 2–4 on a 0–4 point scale for GP5, and ‘severe’ bother was characterised as scores of 3–4. Analyses were conducted separately for each trial.
Results:
Data from 3,557 patients were included. Across the trials, most (71.7%–94.2%) patients had an adverse event of some kind, but fewer (17.1%–44.4%) had an adverse event of grade 3 or higher. In general, fewer than 50% of patients (20.6%–44.2%) reported moderate-severe bother and 5.8%–17.% reported severe bother. There were consistent, albeit not always statistically significant, associations between GP5 and adverse events, and GP5/global health correlations ranged from −0.17 to −0.41.
Discussion:
GP5 is associated with both clinician- and patient-reported symptoms, suggesting its validity and usefulness as part of comprehensive tolerability assessment of cancer trials.
Research article
Available accessResearch articleFirst published April, 2026pp. 166-174
Erica H Brittain, Raphaël N Morsomme, Michael A ProschanORCID
Abstract
Background/Aims:
In randomized two-armed clinical trials with binary endpoints, there may be uncertainty about the event probability, which is needed for sample size calculation. Survival trials are powered based on number of events rather than people, and this is advantageous because the number of events needed to achieve a desired power is less sensitive to an unknown parameter than is the number of people needed. We investigate and quantify this relative stability of number of events compared to number of people in the context of a randomized two-armed trial with equal sample sizes and a binary endpoint. In binary endpoint settings with such relative stability, we consider (1) enhancement of traditional adaptive trial design and (2) potential benefits of a simple event-driven strategy.
Methods:
Using sample size formulas, we determine the relative stability of the expected number of events compared to the sample size for binary outcome trials using the relative risk, odds ratio, or risk difference. Simulations consider a simple event-driven design when there is relative stability; we evaluate type I error rate and power under various analysis methods and approaches to halting the trial.
Results:
We find that the number of events is at least three times more stable than the sample size to achieve a specified power for the relative risk when the overall event probability is less than 1/3, and for the odds ratio when the overall event probability is less than 0.20. We show that this relative stability is independent of the type 1 and type 2 error rates and magnitude of the treatment effect. In a setting where the overall event probability is consistent with relative stability, simulations of an event-driven design show that asymptotic methods may have modestly high type I error rates, but that other approaches appear to have good operating characteristics.
Conclusion:
In settings with moderately low event probabilities, thinking in terms of the number of events instead of sample size may (1) facilitate the planning of clinical trials and help determine whether a trial is futile, and (2) lead to a simple event-driven design for binary endpoints that may be feasible and appealing.
Research article
Open accessResearch articleFirst published April, 2026pp. 175-185
Jordon WimsettORCID, Charlotte Oyston, Robin CroninORCID , [...]
View All
Abstract
Background:
Intrapartum research (occurring during labour and birth) presents challenges to successful recruitment to clinical trials. These include limited time for discussion, decision-making for two (mother and baby), heightened emotional states (pain, anxiety and/or fatigue) and clinician hesitancy to discuss research in this setting. In the context of the Baby head ElevAtion Device Feasibility Study, where the event of interest (caesarean section at full dilatation) is both rare (fewer than 3% of all births) and unpredictable, we undertook a mixed-methods evaluation of the two-stage consent process: (1) abbreviated intrapartum consent and (2) full postpartum consent. The aim was to explore whether abbreviated intrapartum consent was acceptable to patients and clinicians.
Methods:
Eligible patients approached at full cervical dilatation (10 cm) to take part in the Baby head ElevAtion Device Feasibility Study were invited to complete a face-to-face survey of their experience of consent within 3 days after birth. We sampled those who consented and those who declined the study. Clinicians working at recruitment sites were invited to an individual semi-structured interview. Qualitative data were analysed using reflexive thematic analysis.
Results:
Over 12 months, 69% (128/186) of eligible patients consented to the Baby head ElevAtion Device Feasibility Study; 87% of consenters and 66% of decliners completed a follow-up survey. Most survey responders (78%) and clinicians found abbreviated intrapartum consent acceptable. Three themes shaped patient decision-making: perceived benefits, trust in healthcare, and feeling overwhelmed. Those who declined often wished they’d had more time or earlier information. Clinicians found the two-stage consent process feasible and appropriate for low-risk interventions, although time pressures and communication challenges affected consent quality. Many saw the model as respectful of autonomy and potentially useful for future intrapartum research.
Conclusion:
Our findings suggest that the two-stage consent process for this intrapartum study was acceptable to both patients and clinicians. We propose this as a useful consent model for peripartum studies where the clinical situation occurs infrequently, the intervention being studied is low risk, and where opt-out or deferred consent is not available.
Research article
Available accessResearch articleFirst published April, 2026pp. 186-188
The increasing methodological and regulatory demands of clinical trials have increased the need for clearly defined roles, particularly for Clinical Study Coordinators (CSCs) and Data Managers (DMs). While these professionals play crucial roles in the successful conduction of trials, the lack of established consensus on their profile, role, and responsibilities can lead to overlap and inefficiencies. This review aimed to identify the distinct roles of CSCs and DMs, examine their responsibilities, and explore the roles and responsibilities of CSCs and DMs and, when available, aspects of their collaboration.
Methods:
We conducted a systematic review to analyze the distinct roles and responsibilities of CSCs and DMs. A literature search was performed in PubMed, CINAHL, Scopus, and Web of Science for primary studies published between 1 January 2000 and 30 September 2024. Eligible studies focused on defining CSC and DM roles, competencies, and professional responsibilities in clinical trials. Two independent reviewers screened articles, assessed methodological quality, and evaluated the risk of bias. The registration number is PROSPERO CRD42024599819.
Results:
Of the 599 records identified, we included 10 studies. CSCs are responsible for the operational side of trials, including patient recruitment, regulatory compliance, and oversight of trial procedures. At the same time, DMs focus on ensuring data accuracy, integrity, and adherence to regulatory standards. The review highlighted the complementary nature of these roles, suggesting that future research could explore how collaboration between the roles may help maintain data quality and meet the demands of modern clinical research. However, the need for standardized role definitions and formal training programs was a significant challenge.
Discussion:
This systematic review analyzes the roles of CSCs and DMs, emphasizing their complementary responsibilities in clinical trials. It highlights the need for structured training, standardized workflows, and certification programs to enhance efficiency, ensure regulatory compliance, and improve data quality.
Conclusion:
Clarifying the roles of CSCs and DMs and implementing structured training and certification programs are essential to improving trial efficiency. While direct evidence of CSC–DM collaboration was limited, included studies indicate that clear role delineation enhances data quality and regulatory compliance.
Review article
Open accessReview articleFirst published April, 2026pp. 198-209
Kim BoesenORCID, Lars G HemkensORCID, Perrine Janiaud , [...]
View All
Abstract
Introduction:
Conducting systematic reviews of clinical trials is time-consuming and resource-intensive. One potential solution is to design databases that are continuously and automatically populated with clinical trial data from harmonised and structured datasets. This scoping review aimed to identify and map publicly available, continuously updated, topic-specific databases of clinical trials.
Methods:
We systematically searched PubMed, Embase, the preprint servers medRxiv, arXiv, Open Science Framework, and Google. We characterised each database using seven predefined features (access model, database type, data input sources, retrieval methods, data-extraction methods, trial presentation, and export options) and narratively summarised the results.
Results:
We identified 14 continuously updated databases of clinical trials, seven related to COVID-19 (initiated in 2020) and seven non-COVID-19 databases (initiated as early as in 2009). All databases, except one, were publicly funded and accessible without restrictions. Most relied on traditional methods used in static article-based systematic reviews sourcing data from journal publications and trial registries. The COVID-19 databases and some non-COVID-19 databases implemented semi-automated features of data import, which combined automated and manual data curation, whereas the non-COVID-19 databases mainly relied on manual workflows. Most reported information was metadata, such as author names, years of publication, and link to publication or trial registry. Only two databases included trial appraisal information (such as risk of bias assessments). Six databases reported aggregate group-level results, but only one database provided individual participant data on request.
Discussion:
Continuously updated topic-specific databases of clinical trials remain limited in number, and existing initiatives mainly employ traditional static systematic review methodologies. A key barrier to developing truly living platforms is the lack of accessible, machine-readable, and standardised clinical trial data.
Review article
Open accessReview articleFirst published April, 2026pp. 210-218
An estimand is a clear description of the treatment effect a study aims to quantify. The ICH E9(R1) addendum lists five attributes that should be described as part of the estimand definition. However, the addendum was primarily developed for individually randomised trials. Cluster randomised trials, in which groups of individuals are randomised, have additional considerations for defining estimands (e.g. how individuals and clusters are weighted, how cluster-level intercurrent events are handled). However, it is currently unknown if estimands are being used in cluster randomised trials, or whether the considerations specific to cluster randomised trials are being described.
Methods:
We reviewed 73 cluster randomised trials published between October 2023 and January 2024 that were indexed in MEDLINE. For each trial, we assessed whether the estimand for the primary outcome was described, or if not, whether it could be inferred from the statistical methods. We also assessed whether considerations specific to cluster randomised trials were described or inferable, how trials were analysed and whether key assumptions being made in the analysis (e.g. ‘no informative cluster size’) could be identified.
Results:
No trials attempted to describe the estimand for their primary outcome. We were able to infer the five attributes outlined in ICH E9(R1) in only 49% of trials, and when including additional considerations specific to cluster randomised trials, this figure dropped to 21%. Key drivers of this ambiguity were lack of clarity around whether individual- or cluster-average effects were of interest (unclear in 63% of trials), and how cluster-level intercurrent events were handled (unclear in 21% of trials for which this was applicable). Over half of trials used mixed-effects models or generalising estimating equations with an exchangeable correlation structure, which make the assumption that there is no informative cluster size; however, only one of these trials performed sensitivity analyses to evaluate robustness of results to deviations from this assumption. There were 14% of trials that used independence estimating equations or the analysis of cluster-level summaries; however, because no trials stated whether they were targeting the individual- or cluster-average effect, it was impossible to determine whether these methods implemented the appropriate weighting scheme and were thus unbiased.
Conclusion:
The uptake of estimands in published cluster randomised trial articles is low, making it difficult to ascertain which questions were being investigated or whether statistical estimators were appropriate for those questions. This highlights an urgent need to develop guidelines on defining estimands that cover unique aspects of cluster randomised trials to ensure clarity of research questions in these trials.
Research article
Available accessResearch articleFirst published April, 2026pp. 219-224
Pedro A Torres-SaavedraORCID, Boris FreidlinORCID, Jong-Hyeon Jeong , [...]
View All
Abstract
Background:
In randomized trials where some standard-treatment arm patients cross to the experimental treatment, it is frequently of interest to estimate the between-arm survival difference as if no patients on the standard-treatment arm had crossed over to the experimental treatment. Rank-preserving structural failure time models, an extension of semiparametric accelerated-failure-time models, are a popular method for accomplishing this because they do not require modeling which patients will crossover.
Methods:
In trying to apply the rank-preserving structural failure time model in practice, we noted some unusual behavior of the estimated acceleration parameter (differential treatment effect). Simple examples and limited simulations are provided to examine and understand this behavior.
Results:
The simulations show that rank-preserving structural failure time model estimator of the acceleration parameter can take on extreme values, especially when the intent-to-treat analysis favors the standard-treatment arm. Furthermore, the addition of censoring is paradoxically shown to reduce the estimator’s variability compared to the uncensored data when the underlying observations are exponentially distributed. Use of a Weibull distribution with short tails for the survival times eliminates this unusual behavior.
Conclusion:
The rank-preserving structural failure time model estimators of the acceleration parameter are not based on the joint ranks of the original data, and it is suggested that this makes acceleration-parameter estimator unstable with long-tailed survival distributions.
Research article
Open accessResearch articleFirst published April, 2026pp. 225-231
Kylie M LangeORCID, Thomas R SullivanORCID, Jessica KaszaORCID , [...]
View All
Abstract
Background:
Partially clustered trials are trials that, by design, include a mixture of independent and clustered observations. For example, neonatal trials may include infants from a single, twin or triplet birth. The clustering of observations in partially clustered trials should be accounted for when determining the target sample size to avoid treatment arm comparisons being over or under powered. Limited tools are currently available for calculating the sample size for partially clustered trials, particularly when the maximum cluster size is greater than 2. The aim of this article is to introduce a new online application to calculate the target sample size for partially clustered trials covering a broad range of scenarios.
Methods:
The target sample size is calculated using design effects recently derived for two-arm partially clustered trials when the clusters exist prior to randomisation and the outcome of interest is continuous or binary. Both cluster and individual randomisation are considered for the clustered observations (resulting in nested and crossed designs, respectively). The sample size depends on quantities needed for typical sample size calculations, such as the effect size of interest, and the desired significance level and power. In addition, the sample size for partially clustered trials also depends on the range of cluster sizes, the proportion of observations that belong to clusters of each size, the intracluster correlation coefficient, the method of randomisation for the clustered observations, and the model that will be used for analysis. We developed an R Shiny web application that implements these methods in an easy-to-use sample size calculator that is freely available online.
Results:
The sample size calculator is free to access and provides trialists with the ability to determine the target sample size for different types of partially clustered trials. Step-by-step instructions are provided to illustrate the use of the calculator for designing two hypothetical trials. The target sample size that accounts for partial clustering can be quite different to the sample size that is calculated by methods for an independent design that ignore the clustering.
Conclusion:
Partial clustering affects the power and sample size requirements of clinical trials. The calculator presented in this article allows trialists to account for the clustering that occurs in two-arm partially clustered trials for binary and continuous outcomes and ensure their trials are appropriately powered.