We are motivated to understand what influences the extent of hybridization and introgression between xantiana and parviflora. In the previous four chapters, we examined how petal color, petal area, and field site influence the extent of hybridization or introgression. But these variables were not individually modified in isolation. Rather, we put out RILs that differed in numerous traits, in different locations and we observed pollinator visitation. Here, we consider how multiple floral traits and planting location influence the proportion of hybrid seed set on a parviflora RIL.
Treating these separate models as our final analysis is awkward and potentially misleading because:
Biology does not change one variable at a time. Say we were curious about the effect of eating pork on heart attack risk. Simply comparing people who do and do not eat pork would likely drag along many other explanatory variables, because following religious dietary rules of حلال or כַּשְׁרוּת will likely be associated with other practices (e.g., ritualistic prayer) which may decrease stress and increase well-being. So, simply comparing outcomes of people in the pork eating and non-pork eating groups without considering these other variables might miss the actual story. We will see that a similar issue arises in our understanding of traits associated with hybrid seed set in parviflora RILs, below.
Unexplained variation can dampen signal. Although multiple comparisons can give us undeserved confidence in rejecting a null hypothesis, variation that is not accounted for can also decrease our ability to find a real signal. For example, if student scores on standardized tests differ strongly by school, accounting for school can make it easier to detect the effect of participating in a sport. In our example, we do not particularly care about differences in hybrid seed formation by location. But including location in our model can help better estimate the associations between hybrid seed set and the floral traits that matter to us.
Lots of separate tests create a multiple-comparisons problem. If we keep asking many related questions one at a time, some small p-values will appear by chance. Thus, we are risking a false positive at a level greater than our advertised \(\alpha\).
Multiple regression – the focus of this section – addresses these shortcomings by considering simultaneously multiple explanatory variables in one model.
Focal traits for this section
Figure 1: An image of a Clarkia flower to illustrate what “anther stigma distance” means. This is image of a Clarkia xantiana subspecies xantiana flower, as the large anther stigma distance of this species makes this trait easier to see. Figure adapted from and Moeller and Geber 2005.
To introduce work through these ideas, I model hybrid seed set as a function of (some combination of) petal area, petal color, anther-stigma distance (ASD), and planting location.
ASD, the (new to us) variable quantifies the distance between the pollen-producing anther and the ovule-bearing stigma parts of a flower (Figure 1). The logic of this hypothesis is that having male and female parts close to one another can encourage self-fertilization before there is an opportunity for hybridization.
Pairwise associations
In our example, we wanted to know which floral traits were associated with setting more hybrid seed. Individually, each of the three focal explanatory variables is strongly associated with the proportion of hybrid seed set (Figure 2, panels A-C).
But, because biology does not change one variable at a time, any individual pairwise association between one floral trait and hybrid seed set might not mean that this trait itself is directly associated with hybrid seed set. Instead, that trait might be associated with another floral trait, and that other trait might be the one more closely related to hybrid seed set. This holds even in our RIL study. Despite our best efforts to shuffle these traits while constructing our RILs, anther-stigma distance is associated with petal area and petal color (Figure 2, panels D-F).
So, for example, the association between proportion hybrid seed set and petal color (Figure 2, panel A) could arise simply as a consequence of the fact that plants with greater anther-stigma distance are more likely to be pink, and greater anther-stigma distance is associated with lower hybrid seed set, even if petal color itself was unrelated to hybrid seed set. Multiple regression helps separate this knotty tangle of overlapping associations.
With multiple regression, we ask, “Which trait(s) remain associated with proportion hybrid seed set after accounting for the relationships among them?” Or, put another way, “Is petal color associated with hybrid seed set among plants that are similar in petal area and anther-stigma distance, and after accounting for location?”
Figure 2: Proportion of hybrid seed set in parviflora RILs as a function of petal color, petal area, and anther-stigma distance (A-C), as well as relationships among the explanatory variables (D-F). In individual pairwise comparisons, all explanatory variables are associated with proportion hybrid seed. Petal color and petal area are also associated with anther-stigma distance, but are not associated with one another.
Multiple regression as a statistical adjustment
Multiple regression asks this question by introducing a statistical adjustment. Specifically, it aims to estimate the association between one focal explanatory variable and a response variable, after accounting for the other explanatory variables in the model.
In our case, we want to know if e.g. anther-stigma distance is associated with the proportion of hybrid seed set after accounting for petal area, petal color, and planting location. That is, rather than asking whether pink and white flowers differ overall, we ask whether pink and white flowers differ among plants that are similar in these other ways. So we do not just compare all pink flowers to all white flowers. Instead, we compare pink and white flowers while statistically adjusting for other variables that might be tangled up with petal color.
Multiple regression is not a “control”
Often people say they “controlled for” some variable when they mean they included it in their model. This wording can be misleading.
An experimental control is part of a study design. In a controlled experiment, we try to change one variable while holding other relevant variables constant.
A statistical adjustment is part of an analysis. When we include a variable in a regression model, we try to compare observations that are similar with respect to that variable.
Because we did not experimentally manipulate traits independently, multiple regression cannot by itself identify cause.
Adding a “nuisance variable” can sharpen signal
Unless there is something particularly important about our field sites that I don’t know, I do not care if they differ in hybrid seed set. But this does not mean I should not include field site in my linear model.
In this case, location is a “nuisance variable” – something we do not care about, but which may explain variation in our response variable. Including nuisance variables can sharpen the signal associated with the variables we care about most. By including location in the model, we ask whether petal color, petal area, and anther-stigma distance are associated with hybrid seed set among plants from similar locations.
Looking forward
We are now ready to begin our tour of more complex linear models. To understand why this approach may matter, we will focus on one specific question:
Does anther-stigma distance (ASD) still predict the proportion of hybrid seed after we account for the fact that ASD is also associated with petal area and petal color?
Multiple regression is lots of fun! When first learning about it, it’s easy to get carried away and put every variable in the data set into the model. BUT PLEASE DONT. In the best case modeling \(y\) as a function of a grab bag of all explanatory variables that come to mind, generates results that are hard to interpret, and that are statistically under-powered. Whats worse, is that a kitchen sink n of explanatory variables can yield surprising or unstable regression coefficients, and may mistakenly attribute patterns to the wrong variable.
There is an art to building an informative multivariate model. For now, err on the side of “less is more.” We will revisit these ideas later in this chapter and in future chapters.