Abstract
While the IMRAD (Introduction, Methods, Results, and Discussion) format is common in scientific writing, it may not currently be as ubiquitous as often thought. We undertook a systematic, corpus-based study of primary section headings in research articles across a range of STEM disciplines to investigate adherence to the IMRAD structure in relation to type of study (computational, empirical, or theoretical) and field. We identified four categories of structure: IMRAD, IMRAD+ (IMRAD with additional sections and/or different order), Nested IMRAD (multi-part studies), and Non-IMRAD. Papers in biology mainly used an IMRAD format, while less than half in engineering or social sciences did so. While empirical papers tended to use IMRAD formats, most computational papers did not. Thus, our findings show that IMRAD is a common but not universal structure for contemporary scientific writing. Awareness of these differences should encourage teachers of scientific and technical writing and scholars of writing studies to pay closer attention to the actual structural forms used in different STEM disciplines and with different methodological types of research studies.
Keywords
Introduction
The IMRAD (Introduction, Methods, Results, and Discussion) structure for scientific research papers is well-known to researchers and students in every scientific discipline. IMRAD is the required or default structure for scientific articles in many journals as well as for lab reports in most science courses. While one might imagine that this format has long been the norm for scientific writing, it is actually a fairly recent development. The first papers that we would recognize as scientific reports were written in the seventeenth century, but these read more like essays or letters than reports per se (Wu, 2011). Louis Pasteur is credited with writing the first scientific paper resembling modern scientific writing in 1876 (Day, 1989), but the genre did not appear in scientific journals until the 1940s. Leading medical journals began using the IMRAD format in the 1950s, and by the 1960s, IMRAD-structured papers accounted for the majority of medical journal articles. By the 1980s it was the standard form across the health sciences (Sollaci & Pereira, 2004) as well as in some STEM disciplines. Of course, common variations of this structure have developed over time. Gustafson (2011), for example, describes a 7-section variation on the traditional IMRAD structure. (Those interested in the historical development of the IMRAD form might wish to consult Atkinson, 1996; and Bazerman, 1988).
The widespread adoption of IMRAD for scientific writing in the twentieth century is understandable and perhaps inevitable. Meadows (1985) attributes the drastic and sweeping development of the internal organization of research papers in the 1900s to the exponential growth of scientific information. During this period, scientific research and journal editing both came to be performed almost entirely by professional researchers rather than “philosophers of science,” and the consistent, predictable structure of IMRAD had clear advantages for both readers and editors over more essayistic types of writing which had been the norm prior to that point. For one thing, the highly structured format didn’t require scientists to also be capable essayists. More importantly, perhaps, the combination of a standardized format and visually prominent headings made scientific papers much easier to skim. This characteristic became increasingly valuable as the volume of technical papers being published across science and engineering disciplines exploded during the nineteenth and twentieth centuries. The number of scientific journals in existence worldwide grew from less than 100 before 1835 to over 10,000 in 2005. By 2012 there were 28,000 journals which collectively published nearly 2 million articles that year alone (Ware & Mabe, 2012). Estimates of published STEM papers were up to 2.9 million articles in 2020 (White, 2021).
Given the importance of structure in scientific communication, many scholars of scientific writing have conducted investigations on the topic. Of particular note is Charles Bazerman's seminal text Shaping Written Knowledge: The Genre and Activity of the Experimental Article in Science (1988). (See also: Bazerman, 1981, 1984).
A number of corpus-based studies specifically on the meta-structure of research articles have been previously published, but all have been limited to fairly narrow disciplinary realms. Bertin et al. (2013) conducted a corpus analysis of research articles limited to the biological and medical sciences. Heßler et al. (2020) analyzed the structure of original research articles in medical journals. Posteguillo conducted a structural analysis of research articles published in computer science. Graves et al. (2013) studied the organizational structure of research articles in mathematics. A number of cross-disciplinary studies on research article structure have been previously published, but these have primarily investigated rhetorical structure within specific sections rather than the large-scale structure of entire articles. Peacock, for example, analyzed discussion sections in one study (2002) and methods sections in another (2011).
One as-yet uninvestigated factor that might affect adherence to the IMRAD form is the nature of the scientific investigation itself. Research studies in STEM may be roughly categorized into three “study types”: experimental (collecting and analyzing data), computational (using computers to create or refine computational models), and theoretical. Theoretical papers are much less common than the others except in mathematics; in contrast, computational papers have become increasingly common over the past three decades and now represent a substantial fraction of research output across STEM fields—including the quantitative social sciences.
The present study investigates the structure of contemporary research articles across STEM disciplines in relation to the IMRAD model. We sought to determine (1) how common the IMRAD format has been in published articles in recent decades, (2) whether adherence to the strict IMRAD form varies by field, and (3) whether use of the form varied according to “study type” (experimental, computational, or theoretical). We find that IMRAD structure continues to be the dominant format for research articles, but that its use does indeed vary by discipline and study type.
Methods
Data
For this study, we selected a sample of research articles from a previously compiled corpus of published research articles. The articles in that corpus were all produced by research programs supported by U.S. National Science Foundation grants as identified in the papers’ acknowledgement or funding disclosure statements). NSF grants are awarded across a wide range of STEM disciplines including the natural and physical sciences, mathematics, engineering, and the quantitative social sciences. To ensure that the corpus represented the broad range of STEM disciplines, four NSF directorates were chosen—Biology; Engineering; Mathematical and Physical Sciences; and Social, Behavioral, and Economic Sciences—and then two program areas were selected from each of those. The corpus excluded papers from those disciplines whose articles tend to consist primarily of non-prose material (equations or computer coding): mathematics and computer science. Genres other than research articles (e.g., review articles and commentaries) were not included.
A total of 109 standard research grants with grant end dates between 2000 and 2010 were selected such that they were distributed nearly equally among the four disciplinary areas: Biological Sciences (26); Engineering (27); Mathematical & Physical Sciences (29); and Social, Behavioral & Economic Sciences (27). Details about this corpus can be found in Anson & Moskovitz (2021). We then selected a single research article listed for each grant. Each article thus had a single associated field, which we abbreviate as follows: Biological Sciences (BIO); Engineering (ENG); Mathematical & Physical Sciences (MPS); and Social, Behavioral & Economic Sciences (SBE). The papers included in our analysis were published between 1995–2018. A list of all journals included along with the number of papers from each is given in Appendix A.
Analysis of Study Type
Two of us independently categorized each of the papers by study type as either computational, experimental, or theoretical—categories we determined prior to our analysis. We began by examining the article abstracts, searching for keywords that would be indicative of the research category. For computational studies, we looked for descriptions of mathematical modeling (e.g., algorithm, computational, model, simulation). For experimental studies, we looked for descriptions of data collection or experimental methods (e.g., investigation, survey, experiment). We labeled as theoretical those papers which were clearly neither computational nor experimental. When the abstract alone was not sufficient to determine the study type, we examined the major section headings.
Individual categorizations were then compared using Cohen's kappa coefficient. Differences were resolved via re-examination discussion and of the papers in question.
Analysis of Structure
Articles were analyzed to determine their overall structure in relation to the IMRD format on the basis of the main section headings. Four categories were determined via a first pass through the corpus. The corpus was then reanalyzed using those categories:
We counted the number of main sections and determined which sections—regardless of section heading—included descriptions of methodology, recording the title of all sections that did so. We then checked the distribution of these titles across various IMRAD and non-IMRAD format papers.
Results and Discussion
Relationship Between Structure and Study Type
Our calculated kappa statistic for article categorization by study type was 0.605, demonstrating a sufficient level of inter-rate reliability. For 21 of the 109 papers, initial categorized differed and required re-examination.
Table 1 shows the distribution of study types by field: Biological Sciences (BIO); Engineering (ENG); Mathematical & Physical Sciences (MPS); and Social, Behavioral & Economic Sciences (SBE). The highest percentage of experimental papers were in BIO and MPS. The highest percentage of computational articles were in ENG, followed closely by SBE. Very few papers in the corpus were categorized as theoretical.
Distribution of Papers by Field and Study Type: EXP (Experimental), COMP (Computational), THEOR (Theoretical), or Indeterminate.
Table 2 shows article structure as a function of study type. Roughly one third of experimental papers followed the IMRAD or IMRAD + structures; most of the other experimental papers were Non-IMRAD. For computational papers, approximately half were IMRAD or IMRAD+, but only 5 of these followed the strict IMRAD format. None of the four theoretical papers followed an IMRAD or IMRAD + structure. Only six papers were categorized as Nested, and these were all experimental.
A Chi-squared test of the relationships between study type and structure indicated a statistically significant association (7.04, df = 2; p < 0.05).
Article Structure as a Function of Study Type.
Relationship Between Structure and Field
Adherence to the IMRAD structure varied notably by field, as shown in Table 3. BIO papers followed the strict IMRAD format much more frequently than the other fields. Three quarters of MPS papers used either IMRAD or IMRAD +, but less than 14% followed the strict IMRAD format. In contrast, ENG and SBE had a considerable proportion of papers with structures other than IMRAD. In fact, most papers in ENG had non-IMRAD structures. A Chi-squared test of the relationships between field and structure indicated a statistically significant association (51.14, df = 6; p < .05).
Article Structure as a Function of Field.
Variation in Methodology Section Headings
We analyzed the main heading names of sections that provided methodological details for all structure types. Thirteen papers in the corpus lacked a separate methodology section, regardless of section title. A small number of papers had more than one major section containing methodological information: four in ENG, three in SBE, and two in MPS.
We categorized these section headings as either generic (e.g., “Methods,” “Experimental,” “Computational Details”) or content-specific (e.g., “A Nonparametric Bayesian ADCLUS Model,” “A new algorithm for SCSCLP,” “Biochip Fabrication and Surface Functionalization”). The use of generic vs. content-specific methods titles by field is shown in Table 4. Of those papers that followed either the strict IMRAD or IMRAD + format, only 3 used content-specific names for the methodology section. In contrast, the use of content-specific method-section titles was common among papers with other structures. Although our sample of such papers is small, our findings suggest that this practice varies by field: more than half of such papers in ENG (33% of all ENG papers) used content-specific titles for methods sections; for SBE, the proportion was 4 of 11.
Use of Generic vs. Content-Specific Titles for Methods Section by Field and Format.
The most common generic titles for methodology sections are shown in Table 5. The second column presents the raw values along with the proportion of those 96 papers with distinct methodology sections.
Frequently Used Generic Section Titles and Variants for Method Sections.
For non-IMRAD papers, generic names for methods sections varied markedly. Only a single generic methods title was used in more than one non-IMRAD paper: two papers used “Experimental Section.” Other generic method section titles in non-IMRAD format papers were “Computational Model,” “Model,” “Methods,” “Experimental Examples,” “Solution Methodology,” “Computational Experiments,” “Data and Methods,” “Statistical Model,” “First Model (FW1),” and “Second model with set-ups (FW2).” The use of content-specific titles for methods sections in non-IMRAD paper varied by field: ENG, 9/17; MPS, 1/7; and SBE, 4/11.
We also investigated the frequency with which common titles for methods sections (and their minor variants) were used. Table 6 presents the distribution of methods section titles occurring more than once within each field. For two fields, BIO and ENG, one heading was prevalent: “Materials and Methods” in BIO, and variants of “Experimental Section” in MPS. For ENG and SBE papers, there was wider variation.
Distribution of the Most Frequent Generic Method Section Titles (Those Occurring More Than Once in the Corpus) Across Fields. Similar Titles Have Been Grouped for Analysis.
Discussion and Conclusions
While other scholars have investigated the structure of research articles in STEM, we believe ours to be the first study to compare adherence to the IMRAD format across broadly different STEM disciplines and to analyze structure by the study type. Our analysis shows that adherence to IMRAD varies markedly by study type: 32% of experimental papers vs. 19% of computational papers in our corpus closely followed this precise format. And while the majority of experimental papers (nearly 72%) in our sample utilized some variation of IMRAD, many papers—even if generally following the rhetorical sensibilities of IMRAD—use unique, topic-specific section names. The IMRAD format was even less common in computational papers, where nearly half used a different structure and nearly two thirds used topic-specific headings. And even for computational papers that used some variation of IMRAD, 31% used topic-specific rather than generic headings. (While our sample contained only five papers identified as theoretical, none followed the IMRAD or IMRAD variant format.)
Our findings regarding the tendency for computational articles to deviate from IMRAD are in line with the findings of Posteguillo (1999) and Shahid and Afzal (2018). Posteguillo's structural analysis of 40 research articles published in core journals in computer science (in which we can presume that nearly all papers are of a computational nature) could not identify any standard structure for those papers. Shahid & Afzal analyzed a corpus of 329 papers published in the Journal of Universal Computer Science and found none that had a major section titled “Methodology” and that only 1% had a section titled “Results.”
While our study does not provide insight on the reasons why computational papers are less likely to follow the IMRAD model, we can speculate: First, “Methods” and “Results” may be less useful as structural concepts for presenting and explaining new computational models. Second, because computational research is a rapidly developing type of science, authors (and editors) of these papers may have felt less constrained by the standard model developed originally for experimental research. Posteguillo suggested that this lack of standardization in that study might be attributable to the relative youth of the discipline (that study was published in 1999), but our findings suggest that this trend has persisted well into the twenty-first century. Future research might investigate the reasons for these differences through historical and interview studies.
We also found adherence to the strict IMRAD format to be correlated with discipline. Papers in the biological sciences, mathematics, and physical sciences tended toward the IMRAD format and its variants, but in engineering and the social sciences, there was a nearly equal split between IMRAD and non-IMRAD formats. This trend was reflected in the titles for sections containing methodological details: nearly all BIO and MPS papers used generic titles for methods sections, whereas the engineering and the social science papers often used topic-specific titles. It is worth noting that since our corpus uses a single category containing both mathematics and physical sciences, we cannot distinguish between trends in mathematics alone. Graves et al. (2013) found that mathematics papers rarely followed the IMRD form, while we identified 14% of MPS papers as following the strict IMRD structure and another 62% as IMRD + . The trends for mathematics in our study are also affected by the corpus in another way: while a large proportion of papers published in mathematics are deductive (logic-driven) in nature, those that have been supported with NSF grants are more likely to be of an applied nature and thus include more experimental studies.
The generic names used for sections containing methodological information varied by field as well. Most biology papers labeled these sections consistently as “Materials and Methods,” whereas engineering papers used a variety of headings.
Our finding that biological sciences papers were most likely to follow the IMRAD form conflict with those of Bertin et al. (2013). Using a corpus of research articles in the biological and medical sciences from the publisher PLOS, they found that 83% of papers in their corpus contained the four basic sections of the IMRAD format, but the proportion of papers that did so varied markedly across PLOS journals—from a high of over 90% for PLOS One to only 45% for PLOS Biology. The low adherence to IMRAD in biology here is directly contradictory to our results. One possible explanation for this may be that all of the papers in that corpus were published by a single publisher and thus influenced by that publisher’s guidelines.
In contrast to our findings, the study by Heßler et al. (2020) found that the IMRAD structure was used in every paper in their corpus of 450 articles. However, their articles were selected exclusively from medical journals. This exceptional consistency in structure is likely the results of the field of medicine's wide promulgation of discipline-wide publication standards through the “Uniform Requirements for Manuscripts Submitted to Biomedical Journals” starting in 1979 (International Committee of Medical Journal Editors, 1982). Indeed, Sollaci & Pereira's (2004) analysis of medical articles from 1935 to 1985 found that by the end of that period, the IMRAD structure was nearly universal in the medical sciences. Our study—which used a corpus built from papers listed under NSF grants—did not contain medical science papers.
Limitations
The primary limitation is that our sample is drawn from a corpus built exclusively of research articles produced with funding from the U.S. National Science Foundation. This may impact our results in several ways. Most obviously, the principal investigators are likely to be both highly experienced as both scientists and writers of science. Second, papers in the social, behavioral & economic sciences produced with NSF funding are on the quantitative end of the social sciences spectrum and thus certainly not representative of the social sciences generally. Second, our sample did not contain enough papers classified by study type as “theoretical” to draw firm conclusions about such papers. Finally, Of the 109 articles in the corpus, 21 were from journals represented more than once. BIO and SBE each included a single journal represented 5 or more times in the corpus (see Appendix A). Thus, results for these fields may be slightly biased according to the editorial norms for those specific journals.
Because the corpus used did not include the fields of mathematics and computer science, our findings do not apply to those fields.
Practical Implications
Our findings are relevant for future research, teaching, and policy. Our general findings are in line with other studies in showing that the IMRD structure is far from universal and that some of that variation is disciplinarily driven. Scholars of writing and rhetoric should be alert to these differences. But our analysis—which showed different structural norms for computational vs experimental studies—suggests that scholars also need to be careful regarding the date of prior research. Whereas computational studies were still relatively uncommon in the later part of the twentieth century, they have become quite common over the past twenty years. An article published seven years ago in Accounts of Chemical Research (Sperger et al., 2016) established this trend clearly: “Computational chemistry has become an established tool for the study of the origins of chemical phenomena and examination of molecular properties. Because of major advances in theory, hardware and software, calculations of molecular processes can nowadays be done with reasonable accuracy on a time-scale that is competitive or even faster than experiments.” Thus, older studies of rhetorical features in scientific writing in some scientific fields may now be misleading. Peacock's cross-disciplinary study of rhetorical moves in methods sections, for example, includes specific findings for chemistry and physics, among others. However, even though it was published in 2011, the papers in the corpus were published between 2000 and 2003. We might well expect that these findings are no longer representative of those disciplines 20 years later.
Regarding teaching: Many, if not most, teachers of scientific writing and writing center consultants (tutors) are trained in the humanities and have not published any scientific papers themselves. These stakeholders must rely on conventional wisdom—which for the structure of research papers in STEM appears to be oversimplified and, for some disciplines, incorrect. While IMRAD has remained a common format for research writing in STEM disciplines in the twenty-first century, it is sometimes mispresented as universal in statements such as this: “Information within published papers around the world in scientific journals are structured in the format of Introduction, Methodology, Results, and Conclusion (IMRaD)” (Ribeiro et al., 2018). Thus, these teachers and consultants may misinform students about the ubiquity of IMRD papers across the sciences and even suggest specific revisions to drafts in situations where their suggestions may be inappropriate in the discipline for which the student is writing.
Our findings may also rerelevant in the development of scientific writing courses and instructional materials. Most obviously, STEM students should learn that the IMRAD format is common but not universal and be instructed to determine what structural format is expected for the context in which they are working. Implications for instructors are more specific. Those teaching writing in the biological and physical sciences are probably fine to teach IMRAD as the preferred format of their fields, but they should be careful to note that that form is not universal. Those teaching writing in engineering and quantitative social sciences would do well to offer a more nuanced approach, especially so in relation to computational work which has become increasingly common over the past decades.
The policy matter relevant for our findings is in relation to text recycling (AKA “self-plagiarism”)—the reuse of material from one's own previously written documents in a new one. Over the past two decades, a growing number of publishers and journal editors have written or revised their policies regarding text recycling in ways that explicitly tie the acceptability of the practice to the section of a research paper in which the recycling occurs. Consider two examples (emphasis added in all quotations). The first is from BioMed Central/Committee on Publication Ethics (COPE) guidelines regarding text recycling in research articles (BioMed Central, n.d.): Use of similar or identical phrases in methods sections where there are limited ways to describe a method is not unusual; in fact text recycling may be unavoidable when using a technique that the author has described before and it may actually be of value when a technique that is common to a number of papers is described. Editors should use their discretion and knowledge of the field when deciding how much text overlap is acceptable in the methods section. Anesthesiology will permit text recycling (as defined here), when restricted exclusively to a Methods section to describe a standard laboratory method or clinical protocol, and in limited amounts (sentences not multiple paragraphs), and with proper citation to its original publication, and provided it is the author's own prior publication.
Note that both the BioMed Central/COPE guidelines and Anesthesiology policy are framed explicitly in relation to a “methods section.” But as our study shows, many research papers do not have a “Methods” section; instead, methodological details may be in a section with a different heading or they may be scattered over multiple sections that also include results or computational models. The assumption that research papers all follow a strict IMRAD formulation adds unnecessary ambiguity about what is allowed.
In contrast, the Anesthesiology editorial which announced the new policy uses language which is not section-specific (Kharasch et al., 2021): [A]ttitudes and beliefs about limited text recycling … have changed. … In some circumstances…limited text recycling may be appropriate or even preferable. Anesthesiology now provides guidance to authors on how and when to use limited text recycling, specifically limited to describing research methods.
Our findings may have also implications for research in writing studies. For example, in their work on a machine-learning algorithm for categorizing sentences in biomedical articles, Agarwal and Yu (2009) used the four IMRAD sections but did not consider possible variations on these categories or alternative structures. Lu et al. (2018) implemented an algorithm to generate domain-specific functional structure on 13,000 computer science articles. They observed a variety of heading titles but limited their categorization to the IMRAD sections plus an additional one for literature review. Knowing that the structure of research articles in some fields or study types may deviate significantly from the strict IMRAD format may allow for more accurate or nuanced studies.
We conclude but noting that just as IMRAD evolved into the well-known standard of the twentieth century, it will continue to evolve as science and publishing norms change (Wu, 2011). Those who study or teach scientific writing should recognize that the scientific research article is a living rather than a static genre and be alert for changes in discursive norms over time.
Footnotes
Appendix A. Journals included in analysis with number of papers from each
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
