
Editorial
Select search scope: search across all journals or within the current journal

Recent industry–academic partnerships involve collaboration among disciplines, locations, and organizations using publicly funded “open-access” and proprietary commercial data sources. These require the effective integration of chemical and biological information from diverse data sources, which presents key informatics, personnel, and organizational challenges. The BioAssay Research Database (BARD) was conceived to address these challenges and serve as a community-wide resource and intuitive web portal for public-sector chemical-biology data. Its initial focus is to enable scientists to more effectively use the National Institutes of Health Roadmap Molecular Libraries Program (MLP) data generated from the 3-year pilot and 6-year production phases of the Molecular Libraries Probe Production Centers Network (MLPCN), which is currently in its final year. BARD evolves the current data standards through structured assay and result annotations that leverage BioAssay Ontology and other industry-standard ontologies, and a core hierarchy of assay definition terms and data standards defined specifically for small-molecule assay data. We initially focused on migrating the highest-value MLP data into BARD and bringing it up to this new standard. We review the technical and organizational challenges overcome by the interdisciplinary BARD team, veterans of public- and private-sector data-integration projects, who are collaborating to describe (functional specifications), design (technical specifications), and implement this next-generation software solution.
Advances in instrumentation now allow the development of screening assays that are capable of monitoring multiple readouts such as transcript or protein levels, or even multiple parameters derived from images. Such advances in assay technologies highlight the complex nature of biology and disease. Harnessing this complexity requires integration of all the different parameters that can be measured rather than just monitoring a single dimension as is commonly used. Although some of the methods used to combine multiple measurements, such as principal component analysis, are commonly used for microarray analysis, biologists are not yet using many of the tools that have been developed in other fields to address such issues. Visualization of multiparametric data sets is one of the major challenges in this field, and a depiction of the results in a manner that can be readily interpreted is essential. This article describes a number of assay systems being used to generate such data sets en masse, and the methods being applied to their visualization and analysis. We also discuss some of the challenges of applying methods developed in other fields to biology.
Target-based high-throughput screening (HTS) has recently been critiqued for its relatively poor yield compared to phenotypic screening approaches. One type of phenotypic screening, image-based high-content screening (HCS), has been seen as particularly promising.
In this article, we assess whether HCS is as high content as it can be. We analyze HCS publications and find that although the number of HCS experiments published each year continues to grow steadily, the information content lags behind. We find that a majority of high-content screens published so far (60−80%) made use of only one or two image-based features measured from each sample and disregarded the distribution of those features among each cell population. We discuss several potential explanations, focusing on the hypothesis that data analysis traditions are to blame. This includes practical problems related to managing large and multidimensional HCS data sets as well as the adoption of assay quality statistics from HTS to HCS. Both may have led to the simplification or systematic rejection of assays carrying complex and valuable phenotypic information.
We predict that advanced data analysis methods that enable full multiparametric data to be harvested for entire cell populations will enable HCS to finally reach its potential.
Pilot testing of an assay intended for high-throughput screening (HTS) with small compound sets is a necessary but often time-consuming step in the validation of an assay protocol. When the initial testing concentration is less than optimal, this can involve iterative testing at different concentrations to further evaluate the pilot outcome, which can be even more time-consuming. Quantitative HTS (qHTS) enables flexible and rapid collection of assay performance statistics, hits at different concentrations, and concentration-response curves in a single experiment. Here we describe the qHTS process for pilot testing in which eight-point concentration-response curves are produced using an interplate asymmetric dilution protocol in which the first four concentrations are used to represent the range of typical HTS screening concentrations and the last four concentrations are added for robust curve fitting to determine potency/efficacy values. We also describe how these data can be analyzed to predict the frequency of false-positives, false-negatives, hit rates, and confirmation rates for the HTS process as a function of screening concentration. By taking into account the compound pharmacology, this pilot-testing paradigm enables rapid assessment of the assay performance and choosing the optimal concentration for the large-scale HTS in one experiment.
Systematic error is present in all high-throughput screens, lowering measurement accuracy. Because screening occurs at the early stages of research projects, measurement inaccuracy leads to following up inactive features and failing to follow up active features. Current normalization methods take advantage of the fact that most primary-screen features (e.g., compounds) within each plate are inactive, which permits robust estimates of row and column systematic-error effects. Screens that contain a majority of potentially active features pose a more difficult challenge because even the most robust normalization methods will remove at least some of the biological signal. Control plates that contain the same feature in all wells can provide a solution to this problem by providing well-by-well estimates of systematic error, which can then be removed from the treatment plates. We introduce the robust control-plate regression (CPR) method, which uses this approach. CPR’s performance is compared to a high-performing primary-screen normalization method in four experiments. These data were also perturbed to simulate screens with large numbers of active features to further assess CPR’s performance. CPR performs almost as well as the best performing normalization methods with primary screens and outperforms the Z-score and equivalent methods with screens containing a large proportion of active features.
When investigators monitor effects on a population of cells following a perturbation, these events rarely occur in a classical normal (or Gaussian) distribution. A normal distribution is, however, explicitly assumed for events within a single well, in which mean values per well are used as an assay metric and, in general, measures of assay robustness, such as the Z’ score and the V factor. Such analysis is not possible for many technologies; however, high-content screening (HCS) measures events of individual cells, which are averaged over the well. These individual cell-level measurements may be analyzed separately. This study quantifies the extent of nonnormality in experimental samples and their effects on determining the EC50 of a test compound and the assay robustness statistics. The results, based on five sets of publicly available data, indicate that the Z’ or V-factor score can be improved by as much as 0.44 more than standard calculations, and the EC50 of a dose–response curve can be lowered by as much as fivefold when nonparametric methods are used, but not all data sets show a significant improvement. The effect on analysis depends in part on whether the greatest shift from normality occurs in the upper or lower range of the dose–response curve.
High-content screening is a powerful method to discover new drugs and carry out basic biological research. Increasingly, high-content screens have come to rely on supervised machine learning (SML) to perform automatic phenotypic classification as an essential step of the analysis. However, this comes at a cost, namely, the labeled examples required to train the predictive model. Classification performance increases with the number of labeled examples, and because labeling examples demands time from an expert, the training process represents a significant time investment. Active learning strategies attempt to overcome this bottleneck by presenting the most relevant examples to the annotator, thereby achieving high accuracy while minimizing the cost of obtaining labeled data. In this article, we investigate the impact of active learning on single-cell–based phenotype recognition, using data from three large-scale RNA interference high-content screens representing diverse phenotypic profiling problems. We consider several combinations of active learning strategies and popular SML methods. Our results show that active learning significantly reduces the time cost and can be used to reveal the same phenotypic targets identified using SML. We also identify combinations of active learning strategies and SML methods which perform better than others on the phenotypic profiling problems we studied.
A substantial challenge in phenotypic drug discovery is the identification of the molecular targets that govern a phenotypic response of interest. Several experimental strategies are available for this, the so-called target deconvolution process. Most of these approaches exploit the affinity between a small-molecule compound and its putative targets or use large-scale genetic manipulations and profiling. Each of these methods has strengths but also limitations such as bias toward high-affinity interactions or risks from genetic compensation. The use of computational methods for target and mechanism of action identification is a complementary approach that can influence each step of a phenotypic screening campaign. Here, we describe how cheminformatics and bioinformatics are embedded in the process from initial selection of a focused compound library from a large set of historical small-molecule screens through the analysis of screening results. We present a deconvolution method based on enrichment analysis and using known bioactivity data of screened compounds to infer putative targets, pathways, and biological processes that are consistent with the observed phenotypic response. As an example, the approach is applied to a cellular screen aiming at identifying inhibitors of tumor necrosis factor–α production in lipopolysaccharide-stimulated THP-1 cells. In summary, we find that the approach can contribute to solving the often very complex target deconvolution task.
For approximately a decade, biophysical methods have been used to validate positive hits selected from high-throughput screening (HTS) campaigns with the goal to verify binding interactions using label-free assays. By applying label-free readouts, screen artifacts created by compound interference and fluorescence are discovered, enabling further characterization of the hits for their target specificity and selectivity. The use of several biophysical methods to extract this type of high-content information is required to prevent the promotion of false positives to the next level of hit validation and to select the best candidates for further chemical optimization. The typical technologies applied in this arena include dynamic light scattering, turbidometry, resonance waveguide, surface plasmon resonance, differential scanning fluorimetry, mass spectrometry, and others. Each technology can provide different types of information to enable the characterization of the binding interaction. Thus, these technologies can be incorporated in a hit-validation strategy not only according to the profile of chemical matter that is desired by the medicinal chemists, but also in a manner that is in agreement with the target protein’s amenability to the screening format. Here, we present the results of screening strategies using biophysics with the objective to evaluate the approaches, discuss the advantages and challenges, and summarize the benefits in reference to lead discovery. In summary, the biophysics screens presented here demonstrated various hit rates from a list of ~2000 preselected, IC50-validated hits from HTS (an IC50 is the inhibitor concentration at which 50% inhibition of activity is observed). There are several lessons learned from these biophysical screens, which will be discussed in this article.
Although small-molecule drug discovery efforts have focused largely on enzyme, receptor, and ion-channel targets, there has been an increase in such activities to search for protein-protein interaction (PPI) disruptors by applying high-throughout screening (HTS)–compatible protein-binding assays. However, a disadvantage of these assays is that many primary hits are frequent hitters regardless of the PPI being investigated. We have used the AlphaScreen technology to screen four different robust PPI assays each against 25,000 compounds. These activities led to the identification of 137 compounds that demonstrated repeated activity in all PPI assays. These compounds were subsequently evaluated in two AlphaScreen counter assays, leading to classification of compounds that either interfered with the AlphaScreen chemistry (60 compounds) or prevented the binding of the protein His-tag moiety to nickel chelate (Ni2+-NTA) beads of the AlphaScreen detection system (77 compounds). To further triage the 137 frequent hitters, we subsequently confirmed by a time-resolved fluorescence resonance energy transfer assay that most of these compounds were only frequent hitters in AlphaScreen assays. A chemoinformatics analysis of the apparent hits provided details of the compounds that can be flagged as frequent hitters of the AlphaScreen technology, and these data have broad applicability for users of these detection technologies.
High-throughput screening (HTS) is widely used in the pharmaceutical industry to identify novel chemical starting points for drug discovery projects. The current study focuses on the relationship between molecular hit rate in recent in-house HTS and four common molecular descriptors: lipophilicity (ClogP), size (heavy atom count, HEV), fraction of sp3-hybridized carbons (Fsp3), and fraction of molecular framework (
Understanding the structure–activity relationships (SARs) of small molecules is important for developing probes and novel therapeutic agents in chemical biology and drug discovery. Increasingly, multiplexed small-molecule profiling assays allow simultaneous measurement of many biological response parameters for the same compound (e.g., expression levels for many genes or binding constants against many proteins). Although such methods promise to capture SARs with high granularity, few computational methods are available to support SAR analyses of high-dimensional compound activity profiles. Many of these methods are not generally applicable or reduce the activity space to scalar summary statistics before establishing SARs. In this article, we present a versatile computational method that automatically extracts interpretable SAR rules from high-dimensional profiling data. The rules connect chemical structural features of compounds to patterns in their biological activity profiles. We applied our method to data from novel cell-based gene-expression and imaging assays collected on more than 30,000 small molecules. Based on the rules identified for this data set, we prioritized groups of compounds for further study, including a novel set of putative histone deacetylase inhibitors.
In this article, we describe two complementary data-mining approaches used to characterize the GlaxoSmithKline (GSK) natural-products set (NPS) based on information from the high-throughput screening (HTS) databases. Both methods rely on the aggregation and analysis of a large set of single-shot screening data for a number of biological assays, with the goal to reveal natural-product chemical motifs. One of them is an established method based on the data-driven clustering of compounds using a wide range of descriptors,1 whereas the other method partitions and hierarchically clusters the data to identify chemical cores.2,3 Both methods successfully find structural scaffolds that significantly hit different groups of discrete drug targets, compared with their relative frequency of demonstrating inhibitory activity in a large number of screens.
We describe how these methods can be applied to unveil hidden information in large single-shot HTS data sets. Applied prospectively, this type of information could contribute to the design of new chemical templates for drug-target classes and guide synthetic efforts for lead optimization of tractable hits that are based on natural-product chemical motifs.
Relevant findings for 7TM receptors (7TMRs), ion channels, class-7 transferases (protein kinases), hydrolases, and oxidoreductases will be discussed.
Several small-compound library subsets (14,000 to 56,000) have been established to complement screening of a larger Genentech corporate library (~1,300,000). Two validation sets (~1% of the total library) containing compounds representative of the main library were chosen by selection of plates or individual compounds. Use of these subsets guided selection of assay configuration, validated assay reproducibility, and provided estimates of hit rates expected from our full library. A larger diversity subset representing the scaffold diversity of the full library (3.4% of the total) was designed for screening more challenging targets with limited reagent availability or low-throughput assays. Retrospective analysis of this subset showed hit rates similar to those of the main library while recovering a higher proportion of hit scaffolds. Finally, a property-restricted diversity set called the “in-between library” was established to identify ligand-efficient compounds of molecular size between those typically found in fragment and high-throughput screening libraries. It was screened at fivefold higher concentrations than the main library to facilitate identification of less potent yet ligand-efficient compounds. Taken together, this work underscores the value of generating multiple purpose-focused, diversity-based library subsets that are designed using computational approaches coupled with internal screening data analyses to accelerate the lead discovery process.
High-throughput screening allows rapid identification of new candidate compounds for biological probe or drug development. Here, we describe a principled method to generate “assay performance profiles” for individual compounds that can serve as a basis for similarity searches and cluster analyses. Our method overcomes three challenges associated with generating robust assay performance profiles: (1) we transform data, allowing us to build profiles from assays having diverse dynamic ranges and variability; (2) we apply appropriate mathematical principles to handle missing data; and (3) we mitigate the fact that loss-of-signal assay measurements may not distinguish between multiple mechanisms that can lead to certain phenotypes (e.g., cell death). Our method connected compounds with similar mechanisms of action, enabling prediction of new targets and mechanisms both for known bioactives and for compounds emerging from new screens. Furthermore, we used Bayesian modeling of promiscuous compounds to distinguish between broadly bioactive and narrowly bioactive compound communities. Several examples illustrate the utility of our method to support mechanism-of-action studies in probe development and target identification projects.
Small-molecule screens are an integral part of drug discovery. Public domain data in PubChem alone represent more than 158 million measurements, 1.2 million molecules, and 4300 assays. We conducted a global analysis of these data, building a network of assays and connecting the assays if they shared nonpromiscuous active molecules. This network spans both phenotypic and target-based screens, recapitulates known biology, and identifies new polypharmacology. Phenotypic screens are extremely important for drug discovery, contributing to the discovery of a large proportion of new drugs. Connections between phenotypic and biochemical, target-based screens can suggest strategies for repurposing both small-molecule and biologic drugs. For example, a screen for molecules that prevent cell death from a mutated version of superoxide-dismutase is linked with ALOX15. This connection suggests a therapeutic role for ALOX15 inhibitors in amyotrophic lateral sclerosis. An interactive version of the network is available online (http://swami.wustl.edu/flow/assay_network.html).
Gene-expression data are often used to infer pathways regulating transcriptional responses. For example, differentially expressed genes (DEGs) induced by compound treatment can help characterize hits from phenotypic screens, either by correlation with known drug signatures or by pathway enrichment. Pathway enrichment is, however, typically computed with DEGs rather than “upstream” nodes that are potentially causal of “downstream” changes. Here, we present graph-based models to predict causal targets from compound-microarray data. We test several approaches to traversing network topology, and show that a consensus minimum-rank score (SigNet) beat individual methods and could highly rank compound targets among all network nodes. In addition, larger, less canonical networks outperformed linear canonical interactions. Importantly, pathway enrichment using causal nodes rather than DEGs recovers relevant pathways more often. To further validate our approach, we used integrated data sets from the Cancer Genome Atlas to identify driving pathways in triple-negative breast cancer. Critical pathways were uncovered, including the epidermal growth factor receptor 2–phosphatidylinositide 3-kinase–AKT–MAPK growth pathway and
The National Institutes of Health Library of Integrated Network-based Cellular Signatures (LINCS) program is generating extensive multidimensional data sets, including biochemical, genome-wide transcriptional, and phenotypic cellular response signatures to a variety of small-molecule and genetic perturbations with the goal of creating a sustainable, widely applicable, and readily accessible systems biology knowledge resource. Integration and analysis of diverse LINCS data sets depend on the availability of sufficient metadata to describe the assays and screening results and on their syntactic, structural, and semantic consistency. Here we report metadata specifications for the most important molecular and cellular components and recommend them for adoption beyond the LINCS project. We focus on the minimum required information to model LINCS assays and results based on a number of use cases, and we recommend controlled terminologies and ontologies to annotate assays with syntactic consistency and semantic integrity. We also report specifications for a simple annotation format (SAF) to describe assays and screening results based on our metadata specifications with explicit controlled vocabularies. SAF specifically serves to programmatically access and exchange LINCS data as a prerequisite for a distributed information management infrastructure. We applied the metadata specifications to annotate large numbers of LINCS cell lines, proteins, and small molecules. The resources generated and presented here are freely available.
The Bliss independence model is widely used to analyze drug combination data when screening for candidate drug combinations. The method compares the observed combination response (
