Motivating Scenario: You are continuing your exploration of a fresh new dataset. You’ve gotten to know each variable on its own, and now want to summarize associations between two categorical variables.
Learning Goals: By the end of this subchapter, you should be able to:
Calculate and explain conditional proportion: You should be able to do this with basic math and with R code.
Visualize associations between two categorical variables. With pen and paper and R code.
Figure 1: The number of plants that did (light grey) or did not (black) receive a visit from a pollinator, shown as a stacked bar plot.
Figure 2: The percentage of plants that did not (black) or did (light grey) receive a visit from a pollinator, shown as a grouped bar plot.
Even the most complex analysis should begin with clear summaries and visualizations of our data. So, before describing associations between categorical variables, let us revisit our univariate summaries and consider how to summarize categorical variables.
First, we visualize!Figure 1 and Figure 2 show that roughly three quarter of flowers at GC were not visited by a pollinator under our watch. I will show you how to make plots like these below. But for now, I want you to consider the visualization choices available to you, and consider the different strengths and weaknesses of Figure 1 and Figure 2 compared to one another and to other potential visualizations. I will pepper in some guiding data visualization principles as we go, before we address these topics in greater depth at the end of the book. But for now, let these principles guide you:
Show all the data.
Show the data honestly.
Make patterns easy to see.
Next, we summarize! A proportion (e.g., the proportion of flowers receiving a pollinator visit) is a special kind of mean where one outcome (e.g., receiving a pollinator visit) is set to 1 and the other (e.g., not receiving a pollinator visit) is set to 0. So we can calculate a proportion as the number of number of flowers receiving a pollinator visit divided by the total number of flowers with pollinator observations.
We summarize categorical variables in R by taking advantage of the fact that R considers TRUE to equal one and FALSE to equal zero. Within dplyr’s summarise(), we find counts with sum() and proportions with mean():
There are many ways to quantify associations between two categorical variables. Rather than go through them all, I focus here on some key concepts.
Conditional proportions
Figure 3: The association between petal color and pollinator visitation. Petal color is on the x-axis and visit status is shown within bars. We see that pink-flowered plants are more likely to receive a visit from a pollinator.
We found that about three quarters of the RILs planted at GC received no pollinator visits under our watch. But this overall average obscures the key fact – that not all plants are the same. We may believe, and Figure 3 shows, that the proportion of plants receiving visits differs conditional on the petal color! While nearly half of pink-flowered plants received a visit from a pollinator under our watch, only about five percent of white-flowered plants did.
We can represent this in the notation of probability theory as \(P_{A \mid B}\) or P(A|B), meaning the “probability of A conditional on B,” or alternatively, “The probability of A given B”.
P(Visit | Pink) \(\approx\) 0.50.
P(Visit | White) \(\approx\) 0.05.
Below, I introduce janitor’s tabyl() function to aid in these calculations:
If two categorical variables are independent, the conditional probabilities would be equal. It’s clear that, in this case, the probability of us observing a pollinator visit differs conditional on petal color. Pink-flowered RILs had about 10× higher probability (or more technically “relative risk of \(\approx\) 10”) of being visited by a pollinator than did white-flowered RILs.
Additional summaries of associations between categorical variables.
At this point many textbooks would introduce two other standard summaries – odds ratios and relative risk (calculated above). I am not spending much time on them here. That is not because they are not useful (they are) – but because
They can get complicated.
They don’t lead naturally to the next steps in our learning journey.
Feel free to read more about each on Wikipedia (links above) or in conversation with your favorite large language model.
We will also introduce two additional summaries - covariance, and correlation in the next chapter on associations.
Visualizing categorical variables
We have previously introduced making plots with ggplot2. As a quick refresh, let’s consider how to visualize categorical data. Note this section is about how to get a reasonable ggplot up and running - and we will only consider high-level decisions that make plots clear and easy to interpret. Later in the book we will consider how to make better plots.
One categorical variable
The code below shows three ways to visualize one categorical variable: counts, proportions, and side-by-side (aka grouped) comparisons. Each plot (shown in Figure 4) highlights a different aspect of the association. These options can help you choose the visualization that most honestly and clearly displays patterns in your data.
# I am adding this to each plot to clean up the x-axis labelling. Ignore if you likeadded_theme <-theme_light()+theme(axis.text.x =element_blank(), axis.title.x =element_blank(), axis.ticks.x =element_blank(), legend.position ="bottom")# PLOT A ggplot(gc_rils,aes(x =1, fill = visited))+geom_bar()+ added_theme # PLOT B ggplot(gc_rils,aes(x =1, fill = visited))+geom_bar(position ="fill")+labs(y ="proportion")+ added_theme # PLOT C ggplot(gc_rils,aes(x = visited, fill = visited))+geom_bar()
Figure 4: Three ways to visualize visited status of plants.
Two categorical variables
As above, the code below produces three plots (Figure 5) to show the same data in different formats: counts, proportions, and side-by-side (aka grouped) comparisons. Each plot highlights a different aspect of the association.
# I am adding this to each plot to clean up the x-axis labelling. Ignore if you likeadded_theme <-theme_light()+theme(axis.text.x =element_blank(), axis.title.x =element_blank(), axis.ticks.x =element_blank(), legend.position ="bottom")# PLOT A ggplot(gc_rils,aes(x = petal_color, fill = visited))+geom_bar()# PLOT B ggplot(gc_rils,aes(x = petal_color, fill = visited))+geom_bar(position ="fill")+labs(y ="proportion")# PLOT C ggplot(gc_rils,aes(x = petal_color, fill = visited))+geom_bar(position ="dodge")
Figure 5: Three ways to visualize pollinator visits (yes/no) by petal color.