• 6. Two categorical vars

Motivating Scenario: You are continuing your exploration of a fresh new dataset. You’ve gotten to know each variable on its own, and now want to summarize associations between two categorical variables.

Learning Goals: By the end of this subchapter, you should be able to:

  1. Calculate and explain conditional proportion: You should be able to do this with basic math and with R code.

  2. Visualize associations between two categorical variables. With pen and paper and R code.

Alternative formats: 🎥 Watch  ·  🎧 Listen


Loading and processing data.
library(dplyr)
library(readr)
library(ggplot2)

ril_link <- "https://raw.githubusercontent.com/ybrandvain/datasets/refs/heads/master/clarkia_rils.csv"
gc_rils <- readr::read_csv(ril_link) |>
    rename(visits = mean_visits)|>
    mutate(visited = ifelse(visits > 0,"some_visits","no_visits")) |>
    dplyr::filter(location == "GC", !is.na(prop_hybrid), ! is.na(visits),!is.na(petal_color))|>
    select(petal_color, visits, visited, prop_hybrid)

Unconditional proportions

A stacked bar plot showing the number of plants that did (about one quarter) or did not (about three quarters) receive a visit from a pollinator.
Figure 1: The number of plants that did (light grey) or did not (black) receive a visit from a pollinator, shown as a stacked bar plot.
A grouped bar plot showing the percentage of plants that did not (about three quarters) or did (about one quarter) receive a visit from a pollinator.
Figure 2: The percentage of plants that did not (black) or did (light grey) receive a visit from a pollinator, shown as a grouped bar plot.

Even the most complex analysis should begin with clear summaries and visualizations of our data. So, before describing associations between categorical variables, let us revisit our univariate summaries and consider how to summarize categorical variables.

  • First, we visualize! Figure 1 and Figure 2 show that roughly three quarter of flowers at GC were not visited by a pollinator under our watch. I will show you how to make plots like these below. But for now, I want you to consider the visualization choices available to you, and consider the different strengths and weaknesses of Figure 1 and Figure 2 compared to one another and to other potential visualizations. I will pepper in some guiding data visualization principles as we go, before we address these topics in greater depth at the end of the book. But for now, let these principles guide you:
    • Show all the data.
    • Show the data honestly.
    • Make patterns easy to see.

  • Next, we summarize! A proportion (e.g., the proportion of flowers receiving a pollinator visit) is a special kind of mean where one outcome (e.g., receiving a pollinator visit) is set to 1 and the other (e.g., not receiving a pollinator visit) is set to 0. So we can calculate a proportion as the number of number of flowers receiving a pollinator visit divided by the total number of flowers with pollinator observations.
    • We summarize categorical variables in R by taking advantage of the fact that R considers TRUE to equal one and FALSE to equal zero. Within dplyr’s summarise(), we find counts with sum() and proportions with mean():
gc_rils |>
  summarise(n_pink       = sum(petal_color == "pink"),
            n_visits     = sum(visited == "some_visits"),
            n            = n(),
            prop_pink    = mean(petal_color == "pink"),
            prop_visited = mean(visited == "some_visits")) 
n_pink n_visits n prop_pink prop_visited
49 24 91 0.538 0.264

Associations between categorical variables

There are many ways to quantify associations between two categorical variables. Rather than go through them all, I focus here on some key concepts.

Conditional proportions

A bar plot showing the relationship between petal color (pink or white) and pollinator visitation (visited or not visited). Petal color is on the x-axis and proportion of flowers that were visited or not on the y-axis. This plot shows that pink flowers are more likely to be visited.
Figure 3: The association between petal color and pollinator visitation. Petal color is on the x-axis and visit status is shown within bars. We see that pink-flowered plants are more likely to receive a visit from a pollinator.

We found that about three quarters of the RILs planted at GC received no pollinator visits under our watch. But this overall average obscures the key fact – that not all plants are the same. We may believe, and Figure 3 shows, that the proportion of plants receiving visits differs conditional on the petal color! While nearly half of pink-flowered plants received a visit from a pollinator under our watch, only about five percent of white-flowered plants did.

We can represent this in the notation of probability theory as \(P_{A \mid B}\) or P(A|B), meaning the “probability of A conditional on B,” or alternatively, “The probability of A given B”.

  • P(Visit | Pink) \(\approx\) 0.50.
  • P(Visit | White) \(\approx\) 0.05.

Below, I introduce janitor’s tabyl() function to aid in these calculations:

library(janitor)
gc_rils |>
  tabyl(petal_color, visited) 
 petal_color no_visits some_visits
        pink        27          22
       white        40           2

Now that we have these counts, we can use dplyr’s mutate() to find the exact proportions visually estimated from Figure 3.

conditional_table <- gc_rils                                |>
  tabyl(petal_color, visited)                               |> 
  mutate(n_tot =  no_visits + some_visits)                  |>
  mutate(prop_visited =  some_visits / n_tot)      
petal_color no_visits some_visits n_tot prop_visited
pink 27 22 49 0.449
white 40 2 42 0.048

If two categorical variables are independent, the conditional probabilities would be equal. It’s clear that, in this case, the probability of us observing a pollinator visit differs conditional on petal color. Pink-flowered RILs had about 10× higher probability (or more technically “relative risk of \(\approx\) 10”) of being visited by a pollinator than did white-flowered RILs.

Additional summaries of associations between categorical variables.

At this point many textbooks would introduce two other standard summaries – odds ratios and relative risk (calculated above). I am not spending much time on them here. That is not because they are not useful (they are) – but because

  • They can get complicated.
  • They don’t lead naturally to the next steps in our learning journey.

Feel free to read more about each on Wikipedia (links above) or in conversation with your favorite large language model.

We will also introduce two additional summaries - covariance, and correlation in the next chapter on associations.

Visualizing categorical variables

We have previously introduced making plots with ggplot2. As a quick refresh, let’s consider how to visualize categorical data. Note this section is about how to get a reasonable ggplot up and running - and we will only consider high-level decisions that make plots clear and easy to interpret. Later in the book we will consider how to make better plots.

One categorical variable

The code below shows three ways to visualize one categorical variable: counts, proportions, and side-by-side (aka grouped) comparisons. Each plot (shown in Figure 4) highlights a different aspect of the association. These options can help you choose the visualization that most honestly and clearly displays patterns in your data.

# I am adding this to each plot to clean up the x-axis labelling. Ignore if you like
added_theme <- theme_light()+
               theme(axis.text.x = element_blank(),  axis.title.x = element_blank(), 
                     axis.ticks.x = element_blank(), legend.position = "bottom")
# PLOT A 
ggplot(gc_rils,aes(x = 1, fill = visited))+
  geom_bar()+
  added_theme 

# PLOT B 
ggplot(gc_rils,aes(x = 1, fill = visited))+
  geom_bar(position = "fill")+
  labs(y = "proportion")+
  added_theme 

# PLOT C 
ggplot(gc_rils,aes(x = visited, fill = visited))+
  geom_bar()
Three visualization of pollinator visitation. Black bars indicate plants with no pollinator visits; light gray bars indicate plants that received at least one visit. (A) Stacked bar chart showing most plants received no pollinator visits, while a smaller portion received at least one visit. (B) Proportion bar chart showing roughly three quarters of plants had no visits and one quarter had at least one visit. (C) Side-by-side bars comparing counts of plants with no visits versus some visits, showing many more plants without visits.
Figure 4: Three ways to visualize visited status of plants.

Two categorical variables

As above, the code below produces three plots (Figure 5) to show the same data in different formats: counts, proportions, and side-by-side (aka grouped) comparisons. Each plot highlights a different aspect of the association.

# I am adding this to each plot to clean up the x-axis labelling. Ignore if you like
added_theme <- theme_light()+
               theme(axis.text.x = element_blank(),  axis.title.x = element_blank(), 
                     axis.ticks.x = element_blank(), legend.position = "bottom")
# PLOT A 
ggplot(gc_rils,aes(x = petal_color, fill = visited))+
  geom_bar()

# PLOT B 
ggplot(gc_rils,aes(x = petal_color, fill = visited))+
  geom_bar(position = "fill")+
  labs(y = "proportion")

# PLOT C 
ggplot(gc_rils,aes(x = petal_color, fill = visited))+
  geom_bar(position = "dodge")
(A) Stacked bar chart of pink and white flowers showing more visits among pink flowers and very few visits among white flowers. (B) Proportion bar chart showing about half of pink flowers were visited, while only a small fraction of white flowers were visited. (C) Side-by-side bars comparing visit counts by petal color, showing pink flowers received many more visits than white flowers.
Figure 5: Three ways to visualize pollinator visits (yes/no) by petal color.