Abstract
Many academic researchers regard logistic regression as the preeminent analytic approach for modeling binary outcomes. It can identify and estimate the effects of actions to increase or decrease the size or proportion of the group of interest. It can also predict each case’s probability of belonging to one group instead of another, given the model’s explanatory variables. However, evidence indicates that market researchers do not use it extensively to analyze survey data, partly because of the difficulty in translating logistic regression’s standard analysis output—logits, odds, and odds ratios—into clear, action-oriented findings and recommendations. The aim here is to offer an informed view, supported by analysis of Pew Research Center survey data, of the possible benefits of reporting percentage point effects (e.g., a one-unit change in x is associated with a three-percentage-point increase in y, all else being equal), in addition to logits, odds, and odds ratios. Such reporting may help to reduce any gap between what some clients expect—particularly when they ask researchers to identify and estimate the effects of actions for increasing or decreasing a critical group’s size or proportion—and what they may receive in return. It may also create new consulting and relationship-building opportunities for market researchers.
Keywords
Overview
Logistic regression models the relationship between a binary 1 outcome (e.g., customer or non-customer, or nearly anything with a yes or no interpretation) and, typically, several explanatory variables. 2 It can identify and estimate the effects of actions to increase or decrease the size or proportion of the group of paramount interest. It can also predict each case’s probability of belonging to one group rather than another.
Many academic researchers consider it “the standard way to model binary outcomes” (Gelman & Hill, 2009, p. 79), possibly “dominating all other methods in both the social and biomedical sciences” (Allison, 2015). However, evidence indicates that market researchers do not use it extensively to analyze survey data, despite a client need across service lines (e.g., customer experience monitoring, brand health monitoring, concept testing, advertising testing, political polling) to understand how two groups differ, often a necessary step toward identifying effective actions for increasing or decreasing a key group’s size or proportion. The evidence includes reviews of journal articles, 3 conference papers, and presentations and personal communication with more than 125 current or former employees 4 (mainly, marketing scientists, data scientists, and methodologists but also chief executive officers, salespeople, and others) from 11 of the 15 largest global market research agencies. 5
The evidence suggests that difficulty in translating logistic regression’s standard analysis output—logits, odds, and odds ratios—into clear, action-oriented findings and recommendations is the main reason for its inextensive use. 6 A different way to say this is that some market research clients apparently have struggled to interpret and act on findings and recommendations communicated in logits, odds, and odds ratios, particularly when they posed their initial research questions in proportions (e.g., “What actions should we consider for increasing the proportion of Americans who are enthusiastic about driverless vehicle development? By how many percentage points would we expect each action to increase that proportion, controlling for other variables’ effects?”).
Although it is possible to report the effect of an explanatory variable, x, on a binary outcome, y, in percentage points (e.g., a one-unit change in x is associated with a three-percentage-point increase in y, all else being equal), the size of the effect will depend both on the value of y and on the values of the model’s other explanatory variables. As a result, x’s effect on y in percentage points “. . . cannot be fully represented by a single number” (Pampel, 2000, p. 23). This may be why some logistic regression experts (e.g., DeMaris, 1990, 1992) have advised against using percentage points to interpret and report logistic regression coefficients’ overall effects. It may also be why most major statistical software packages do not produce percentage point effects through prepackaged procedures or built-in modules.
The aim here is to offer an informed view of the possible benefits of reporting percentage point effects, in addition to logits, odds, and odds ratios. The general idea, to borrow from the statistician Frederick Mosteller (1996), would be to let “weaknesses from one method . . . be buttressed by strength from another” (Ch. 4, p. 116), a concept he referred to as “balancing biases.” 7 Such reporting may help to reduce any gap between what some clients expect—particularly when they ask researchers to identify and estimate the effects of actions for increasing or decreasing a critical group’s size or proportion—and what they may receive in return.
This article has six sections. The first section explains why linear regression may not be fit for the purpose when the outcome of interest is binary rather than continuous. The second section, using Pew Research Center survey data on Americans’ views about driverless vehicles for illustrations and examples, describes the mathematical concepts underlying logistic regression analysis. The third section reviews the main options for interpreting and reporting the effects of logistic regression coefficients. The fourth section explains how to calculate those effects in percentage points, while the fifth section builds on earlier analyses of Pew data to show how percentage point effects reporting could complement logit, odds, and odds ratio reporting. The last section considers the implications of the ideas, suggestions, and new research presented here.
The challenge of modeling binary outcomes through linear regression
“What effect does x have on y?” can be a key question for market research. Many methods can help to answer the question, including randomized controlled experiments, statistical matching, and several types of regression analysis. The data’s character can also influence the decision on which method to apply. If the critical outcome variable is continuous, and a controlled experiment is not feasible, then linear regression might be the right choice. A linear regression model might show that a one-unit change in x is associated with a ten-unit change in y, all else being equal. To produce this estimate, it would find the best straight-line predicting y from x using ordinary least squares estimation.
The x, y relationship can be expressed by the equation y = a + bx, where y is the outcome, a is a constant and y’s value when x equals 0, and b is the slope or the change in y associated with a one-unit change in x. With multiple explanatory variables, the equation can be extended by adding x’s (e.g., x2, x3) and b’s (e.g., b2, b3), or y = a + b1x1 + b2x2 + . . . bNxN.
If the outcome variable is binary, then linear regression may not produce credible, trustworthy information. A model could predict that some outcome probabilities are negative while others exceed 1, even though the scale is bounded by 0 and 1. When valid predictions are essential, this can lead to awkward, uncomfortable moments. As Allison (2017) remarked, “. . . if you want to give osteoporosis patients an estimate of their probability of hip fracture in the next five years, you won’t want to tell them it’s 1.05.” 8
The problem is that x’s effect on y becomes compressed near 0 and 1 on the probability scale. So, trying to use a straight line 9 to predict y from x may not work well. One way to address this and related issues (e.g., unstable b’s) is by swapping out linear regression’s straight line for a curve that runs from negative to positive infinity. The idea behind the curve, according to Pampel (2000), is to stretch or extend probabilities near 0 and 1 so that “the same change in x comes to have similar effects” (p. 15) for all predicted y values. He referred to this as “linearizing the nonlinear” (p. 14) x, y relationship.
To better understand the approach, which is logistic regression’s foundation, some knowledge of probabilities, odds, odds ratios, and logits can be helpful because several transformations—probabilities to odds, odds to odds ratios, odds ratios to logits—take place to make the underlying math work.
Making logistic regression’s math work
The following examples and illustrations rely on Pew Research Center data, collected online through a survey of 4,135 US adults in May 2017. A report titled, “Automation in Everyday Life” (Pew Research Center, 2017) contains the main findings, commentary, and other methodological details.
Table 1 shows that 40% of US adults 10 say they are enthusiastic about driverless vehicle development, with men more enthusiastic than women: 46% versus 34%. Each percentage can be thought of as a probability.
Enthusiasm of US adults about driverless vehicle development.
Odds represent the ratio of a probability (e.g., the probability, p, of being male) to its non-probability, or p/(1 − p). Men’s odds of being enthusiastic about driverless vehicle development are .85, or .46/(1 − .46); women’s odds are .51, or .34/(1 − .34).
The ratio of one to the other indicates relative enthusiasm about driverless vehicle development. The male-to-female odds ratio is .85/.51, or 1.67 (to 1); the female-to-male odds ratio is .51/.85, or .6 (to 1).
The relationships can be described through multiplication where men’s odds are .85 = .85 * 1 and women’s odds are .51 = .85 * .6 (alternatively, women’s odds of .51 = .51 * 1 and men’s odds of .85 = .51 * 1.67). These numbers suggest each gender’s odds can be thought of as the product of a constant and a gender-specific factor: the odds ratio. By replacing the constant with the letter “a,” the result is the equation p/(1 − p) = a * the odds ratio.
The logit, ln, or the natural logarithm of the odds is the power to which e, or the (approximate and rounded to the fourth digit) “irrational” number 2.718, must be raised to equal the odds. Men’s logit of being enthusiastic about driverless vehicle development is −.16, or ln(.85). Put differently, −.16 is the answer to the question, “To what power must we raise 2.718 to equal .85?” Women’s logit is −.67, or ln(.51).
A feature of logits is that they transform multiplication and division to addition and subtraction. Accordingly, the odds ratio in logits for women to men is −.51, or −.67 − −.16 and the male to female odds ratio in logits is .51, or −.16 − −.67. Researchers can interpret the .51 absolute difference as the change in logits in y associated with a one-unit change in x, as in linear regression.
Given this feature of logits, the x, y relationship can be expressed through the equation where the logit of y, or ln, p/(p − 1), = the logit of a constant (a) + the logit of the odds ratio (b). 11 Inserting the letter “x” after b, or ln(p/1 − p) = ln(a) + ln(bx), provides a way to distinguish between men and women.
The equation can be extended to include more explanatory variables: ln(p/1 − p) = ln(a) + ln(b1x1) + ln(b2x2) + . . . ln(bNxN). It may look familiar because it is the linear regression equation shown earlier, except in logits. In other words, the nonlinear relationship between x and y has been linearized.
Rather than estimating these b’s through least squares, as in linear regression, it is considered the best practice to use maximum likelihood in logistic regression. The procedure begins by assigning arbitrary estimates, or starting values, to each b. It then adjusts these values iteratively to maximize their joint effectiveness at predicting the actual y’s correctly.
Through these steps, it is then possible to estimate the effect on a binary y of one or more x’s via a model that is linear in logits. In the multiple explanatory variable (or “multiple x”) model, however, it is more difficult than in the “single x” model (e.g., when “gender” was the lone explanatory variable) to interpret (each) x’s effect on y in percentage points. As noted earlier, a constant effect in logits often translates into a nonconstant effect in percentage points.
To show how this works, Table 2 lists the illustrative values of logits, their associated probabilities, and, to complete the picture, their corresponding odds. Note how logits are symmetrical around 0 and run from negative to positive infinity, probabilities are bounded by 0 and 1, while odds have a floor at 0 but no ceiling—they increase by multiples of 2.718 as logits increase by 1. A four-unit logit increase from 1 to 5, for instance, would translate to a 2.7184 odds increase of 54.6%, or 148.4/2.7.
Illustrative values of logits, probabilities, and odds.
Options for interpreting and reporting explanatory variables’ effects
As the above paragraphs point out, a benefit of what some researchers call the “logit transformation” is linearization of the nonlinear x, y relationship. Logistic regression, through this lens, can be thought of as an enhancement of linear regression for binary outcome variables. But it is more challenging in logistic than linear regression to interpret each explanatory variable’s effect.
Traditionally, researchers have relied on some combination of logits, odds, and odds ratios to do so. These measures have merits, but ease of interpretation and actionability may not top the list. Consider the statement: “A one-unit (or one-category) change in gender (i.e., from female to male) increases the logit of being enthusiastic about driverless vehicle development by .51.” Or “men’s logit of being enthusiastic about driverless vehicle development is .51 higher than women’s.” Without more information, what these statements mean is unclear. As a reminder, a logit is an exponent, not the usual type of number on which market research clients rely.
A second option is to convert logit coefficients to odds ratios through exponentiation, or by raising e to the applicable logit power. Hearing men’s odds of being enthusiastic about driverless vehicle development are 67% or 1.67 (i.e., 2.718.51) times higher than women’s may be easier to grasp than a logit-only statement. That odds have no ceiling can be appealing, too, especially when a research goal is to identify important x’s irrespective of their percentage point effects on y. As Allison (2017) explained, “If the probability that I will vote in the next presidential election is .6, there’s no way that your probability can be twice as great as mine. But your odds of voting can easily be 2, 4, or 10 times as great . . .”
Odds and odds ratios do have critics, including Gelman and Hill (2009), who asserted, “. . . odds can be somewhat difficult to understand, and odds ratios are even more obscure” (p. 83). From a client’s perspective, moreover, odds and odds ratios do not answer the question, “By how many percentage points would we expect each action to increase the proportion of interest, controlling for other variables’ effects?”
Percentage point effects reporting, a third, less-conventional option, answers that critical client question. It can also promote return-on-investment (ROI) analysis, as a later example shows. But converting logit coefficients to percentage points, as noted earlier, can create interpretive challenges in models with more than one (categorical) explanatory variable because of the nonlinear relationship between logits and probabilities. To reinforce this point visually, Figure 1 plots the illustrative logit and probability values shown in Table 2.

The relationship between logits and probabilities.
Note how a one-logit increase from 0 to 1 on the x-axis corresponds to a .23 probability increase (from .5 to .73) on the y-axis. Yet a one-logit increase from 5 to 6 (or from −6 to −5) translates only to a minuscule probability increase. DeMaris (1993) considered this (i.e., how a constant effect in logits can turn into a non-constant effect in probabilities) as an “intractable” (p. 1,057) problem and sufficient reason to use logits, odds, and odds ratios when interpreting and reporting explanatory variables’ overall effects.
For a market research client needing to learn how to increase or decrease a crucial proportion, however, it may reflect reality. Consider, for example, a company investing in innovative automation technology to support driverless vehicle development. It could launch a social media campaign to raise Americans’ enthusiasm for driverless vehicles. Although the campaign may appeal to like-minded driverless vehicle proponents, it may do little to raise their already high probability of being enthusiastic about driverless vehicle development. It may also do little to increase the probability of those at the spectrum’s other end—Americans who would rather be barricaded in their homes than on the road with driverless vehicles—to transform near-immediately into proponents. The campaign probably would make more of an impact on Americans in the middle as Figure 1’s elongated s-shaped curve would suggest.
For clients believing x’s effect on y in percentage points should be smaller near 0 or 1 than .5 on the probability scale, a question would remain on how to calculate this effect.
Calculating percentage point effects
Logistic regression generates for each case (e.g., a Pew survey respondent referred to here, for convenience, as “Morgan”) a predicted probability of belonging to the group of interest. The simple model shown in Table 3 indicates, for instance, that Morgan’s predicted probability of being enthusiastic (vs non-enthusiastic) about driverless vehicle development is .95 (or 2.90 in logits), 12 given her characteristics. She is 35, earns US$174,000 a year, lives just outside Las Vegas, Nevada, would feel very safe on the road with driverless vehicles, and believes driverless vehicles’ widespread use would lead to much less traffic in major cities. 13 The equation, in logits, would look like this: Morgan’s predicted probability of (2.90) = constant (2.63) + age (−.06) + gender (.11) + household income (0) + region (.22) + feel safe? (0) + less traffic? (0).
Results of logistic regression analysis (simple model).
n = 4,028.
The 40% base value in the far-right column refers to Americans who say they are enthusiastic about driverless vehicle development.
Log pseudolikelihood, starting value: −2722.8634; final value: −1867.0644.
Wald chi (13): 380.52; Prob > chi2: .00.
Stukel goodness of fit: chi2(2) = 2.55; Prob > chi2 = .2798.
McFadden R2: .32; Tjur R2: 38.
Data were weighted using the variable weight_W27.
Each parenthetical number, excluding the one (i.e., 2.63) to the constant’s immediate right, is the logit coefficient corresponding to Morgan’s associated attribute (e.g., .11 is the coefficient for female). For the constant, the number is the sum of the logit coefficients for the reference categories: age 18–29, male, household income of US$75,000 or higher, lives in the Northeast, would feel very safe on the road with driverless vehicles, believes driverless vehicles’ widespread use would lead to much less traffic in major cities.
After reviewing this information, a research client may wonder how Morgan’s .95 probability would have changed if she instead believed that driverless vehicles’ widespread use would
To respond, the researcher could replace her less traffic? coefficient of 0 with −.82, the one corresponding to a No, not likely answer. It would reduce Morgan’s summed logit score from 2.90 to 2.08, and her predicted probability from .95 to .89. A one-unit change in x, therefore, would result in a .06 decrease in y, all else being equal, with .06 (or six percentage points) the percentage point effect. 14
The client then might ask the researcher to estimate the percentage point effect of a one-unit change in the less traffic? variable for the entire Pew sample. As context, 28% of the sample responded Yes, likely while 72% responded No, not likely when asked if they thought the widespread use of driverless vehicles would lead to “much less traffic” in major cities.
To address this request, the researcher could change the value of the less traffic? binary variable to the one each respondent did not choose, 15 calculate a new predicted probability, then take the difference between the original and the new. 16 The mean of these differences across all respondents, or .13 (i.e., .50 − .37), would be the percentage point effect on y of a one-unit change in the less traffic? variable.
The researcher then could share the following information with the client: “All else unchanged, if all Americans, rather than 28%, thought driverless vehicles’ widespread use would lead to much less traffic in major cities, then the percentage of Americans who say they are enthusiastic about driverless vehicle development would increase from 40% to 50%. But if all Americans, rather than 72%, thought it would
A point to note is that the less traffic? variable’s effect is about two times larger for all Americans than for Morgan (i.e., 13 vs 6 percentage points), primarily because her probability of being enthusiastic about driverless vehicle development was quite high already. As described earlier, the size of an explanatory variable’s effect in percentage points depends both on the value of y and on the values of the model’s other explanatory variables. As Figure 1 shows, the size of the effect is smaller near the probability scale’s ceiling and floor than its middle.
To estimate the effect on y in percentage points of a one-unit change in the value of any other explanatory variable, or the effect of simultaneous one-unit changes in the values of two or more variables, the researcher could carry out this same procedure. 17 Through an experimenter’s eyes, it would be analogous to conducting one or more post hoc 18 simulated quasi-experiments.
A deeper dive into pew research center data
Through added, more-comprehensive analysis of Pew data, this section aims to show how reporting percentage point effects might complement logit, odds, and odds ratio reporting.
Pew commented that “Most Americans are aware of the effort to develop driverless vehicles and express somewhat more worry than enthusiasm about their widespread adoption” (p. 29).
Pew also noted Americans strongly favor policies such as “requiring driverless vehicles to travel in dedicated lanes” (p. 36) and “restricting them from traveling near certain areas, such as schools” (p. 36).
But Pew did not estimate, through statistical modeling, the effect of those or other variables on Americans’ enthusiasm about driverless vehicle development. Consequently, the report does not answer a critical strategic question a driverless vehicle developer may ask: “By how many percentage points would we expect each action to increase the proportion of Americans who are enthusiastic about driverless vehicle development, controlling for other variables’ effects?”
For the new analysis, the outcome variable is the same as that used earlier: whether Americans say they are enthusiastic about driverless vehicle development. 19 The explanatory variables, all categorical, include several socio-demographic and opinion-based ones. They were selected based on their univariate relationship with the outcome variable and one another, theory, and availability.
Table 4 contains the logistic regression analysis’s results, including standard information such as logit coefficients, odds ratios, z scores, and the McFadden R2. It also includes nonstandard information such as predicted probabilities, percentage point effects, and the Tjur R2.
Results of logistic regression analysis (full model).
n = 3,748.
The 42% base value in the far-right column refers to Americans who say they are enthusiastic about driverless vehicle development.
Log pseudolikelihood, starting value: −2503.30; final value: −1551.09.
Wald chi (27): 478.99; Prob > chi2: .00.
Stukel goodness of fit: chi2(2) = 0.44; Prob > chi2 = .8036.
McFadden R2: .38; Tjur R2: .44.
Data were weighted using the variable weight_W27.
The explanatory variable with the largest effect is the response to the question, “How safe would you feel sharing the road with a driverless passenger vehicle?” A typical interpretation would emphasize odds, odds ratios, and statistical significance. It would read like this: “Controlling for other variables’ effects, Americans who say they would feel ‘very safe’ sharing the road with a driverless vehicle have a 69% higher odds of saying they are enthusiastic about driverless vehicle development than those who say they would feel ‘somewhat safe,’ a 91% higher odds than those who say they would feel ‘not too safe,’ and a 98% higher odds than those who say they would feel ‘not safe at all.’ Each effect is statistically significant, as their z scores show.”
Although the interpretation is correct, a client or other interested party may find it challenging to act on because it does not show how an increase in the percentage of Americans who say they would feel “very safe” would change the enthusiastic group’s size or proportion.
Now consider an alternative interpretation: “All else unchanged, the percentage of Americans who say they are enthusiastic about driverless vehicle development would increase from 42%, 20 the current level, to 72%, the new level, if all Americans were to say they would feel ’very safe‘ sharing the road with a driverless vehicle. At the other extreme, if all were to say they would feel ’not safe at all,’ that same percentage, 42%, would drop to 9%.”
Some research clients may prefer the alternative interpretation because it reports the effect of the explanatory variable in percentage points, often a more action-oriented measure than odds.
As a second example, consider how a researcher might interpret the effect of gender on the outcome. A standard interpretation would highlight the statistically significant finding that females’ odds are 38% higher than males’ of saying they are enthusiastic about driverless vehicle development, controlling for other variables’ effects. But it would leave open the question of how the gender difference translates to probabilities.
Percentage point effects reporting would answer the question, showing females have a four-percentage-point higher predicted probability than males, 44% versus 40%, of being enthusiastic about driverless vehicle development, all else unchanged. Although the difference may not be earth-shattering, it is telling because “men are a bit more likely than women [46% vs 34% as Table 1 shows] to say they are enthusiastic about driverless vehicle development” (Pew Research Center, p. 30). When more explanatory variables are added to the model, however, the relationship turns on its head: women have a higher predicted probability than men. An implication is that women could become stronger supporters than men of driverless vehicle development, especially if their concerns about safety are allayed.
To increase the usefulness of these and other findings from the logistic regression analysis, a company developing driverless vehicles, or some other interested party, could reorganize the information in Table 4 by sorting all predicted probabilities in descending order, as shown in Figure 2. 21

Predicted probability of each variable in descending order.
After seeing the .72 predicted probability associated with the feel very safe? response choice, the driverless vehicle developer might decide to make “safety” a focal point of future advertising campaigns, perhaps believing its eventual business success would hinge partly on increasing Americans’ safety perceptions toward driverless vehicles. It might set campaign goals and measure a type of ROI later by depending on evidence from the Pew survey and percentage point effects reporting. Through these sources, it would know the following:
In all, 42% 22 of Americans say they are enthusiastic about driverless vehicle development.
While 11% of Americans say they would feel very safe sharing the road with a driverless vehicle.
If all Americans were to say they would feel very safe sharing the road with a driverless vehicle, then the percentage of Americans who say they are enthusiastic about driverless vehicle development would increase to 72%, all else being unchanged. An 89-percentage-point increase on the feel very safe? explanatory variable, therefore, would be associated with a 30-percentage-point increase on the outcome variable, a three-to-one ratio.
Accounting for this and other information, the developer then might invest an incremental US$10 million in advertising in the next year with a goal of more than doubling, from 11% to 23%, the percentage of Americans who say they would feel very safe sharing the road with a driverless vehicle. Given the three-to-one ratio, the developer could expect the 12-point lift on the feel very safe? response choice to increase the percentage of Americans who say they are enthusiastic about driverless vehicle development by about four points, from 42% to 46%. 23
Estimating the advertising’s ROI at year’s end would call for basic math. 24 If the advertising achieved its goals, the cost per-percentage-point increase on the feel very safe? explanatory variable would be US$833,333, or US$10 million/12, while the cost per-percentage-point increase on the enthusiastic? outcome variable would be US$2.5 million, or US$10 million/4. (There are about 250 million American adults, so the cost per-adult increase would be US$0.33 and US$1, respectively.)
Had the Pew report included percentage point effects reporting driverless vehicle developers, including companies such as Google, Waymo, Apple, BMW, Tesla, and Baidu may have found it even more illuminating.
Closing remarks
The suggestions, ideas, and evidence presented here suggest that market researchers, at times, may be able to produce clearer, more-action-oriented findings and recommendations through logistic regression by using percentage points to report explanatory variables’ effects. Clients with a keen interest in learning how to increase or decrease a key proportion may welcome the capability. It may also create new consulting and relationship-building opportunities for market researchers. Given that binary outcomes are common, they may have many opportunities to report percentage point effects:
Customer Experience Monitoring: one-time versus repeat customer, promoter versus detractor
Brand Health Monitoring: love the brand or not, use the brand or not
Concept Testing: likely to consider purchasing the product or not, willing to pay a high price or not
Advertising Testing: love the ad or not, click through the ad or not
Political Polling: voter versus non-voter, support versus oppose the policy
To exploit percentage point effects reporting’s full benefits, market researchers may also want to consider modifying the design of selected surveys and other information systems. Conceptualizing them as platforms for estimating (potentially causal) effects could be a good first step. 25 A second step might involve ensuring these newly designated “platforms” include the necessary variables to permit and promote rigorous post hoc quasi-experimentation. In principle, the y’s should be true dichotomies and the x’s, aside from control variables (e.g., socio-demographic questions), should be action-oriented levers clients can pull to affect the size of a key proportion. 26 The rationale for drawing attention to research design is straightforward: applying an analytical method, no matter how promising, to data from a survey or other information system not designed with that method in mind may bear little fruit.
An overemphasis on percentage point effects, however, could cause some unwary researchers, whether or not they are analyzing survey data, to overlook a model’s important explanatory variables because of the nonlinear relationship between logits and probabilities. In online advertising, for example, click-through rates at times are below 1%. If click-through were the outcome variable in a logistic regression model, then the percentage point effect of a one-unit change in an important explanatory variable could fall under a researcher’s radar.
An analysis might suggest, for example, that the use of active-voice language in call-to-action display ads, controlling for other variables’ effects, increases the click-through probability from .0025 to .0067, a mere fraction of a percentage point and possibly easy to overlook. Standard reporting would supply the information (e.g., logits, z scores, odds ratios) needed to reduce the risk of an oversight: the .0042 percentage point increase would translate to a substantial 172% odds, and a one-point logit, increase.
It is a good example of Mosteller’s “balancing biases” concept of letting “weaknesses from one method . . . be buttressed by strength from another.” But in this case, standard reporting would offset a possible shortfall (i.e., tiny effects near 0 on the probability scale) of percentage point effects reporting, rather than the other way around. An implication is that the two approaches can work hand in hand.
Footnotes
Appendix
Acknowledgements
The author would like to thank Reg Baker, Regina Corso, Robert Eisinger, Ward Fonrose, Ryan Heaton, Sylvie Lacassagne, Scott McDonald, Mark Naples, Terry Sullivan, Jennifer Timko, and two anonymous reviewers for helpful suggestions on earlier versions of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
