• 12. Threats to Validity
Motivating Scenario:
You want to make sure your study has a reasonable chance to answer the question you care about - not some other question.
Learning Goals: By the end of this chapter, you should be able to:
- Recognize what makes a study “valid”.
- Distinguish measurement validity from measurement reliability.
- Know what “internal validity” means, and recognize what can limit the internal validity of our study.
- Know what “external validity” means, and describe what limits how far a study’s results generalize beyond the conditions it was run under.
- Explain why every design decision is a compromise, why we should not try to have a perfect design, and why “good enough to inform the question” is good enough.
Note: Your study is not the last word on the topic. At best it meaningfully informs your scientific question. Remember science builds on itself, and thrives on replication.
Designing a scientific study requires us to make many choices. Nearly all these choices reflect some compromise between the ideal study and the one we can actually pull off. Time, money, and feasibility will naturally limit the science we can do. But it is our job to make decisions that best set us up for a study that will inform our scientific question.
Here, I discuss some of the decisions we made in this study and how they relate to various aspects of the “validity” of a scientific study.
Measurement validity and reliability
We decided to measure pollinator visitation and the proportion of hybrid seeds on a RIL (out of a sample of eight). These two measures bring up ideas of measurement:
Measurement validity (aka Construct validity) describes how well the thing we measure connects to the thing we care about. In this case, for example, we care about hybridization. So pollinator visitation may be an ok proxy. But it is not the thing we care about. Similar challenges occur across biology. For example, we may want to measure “organismal fitness” or “health” or “understanding”, but we cannot always directly access these concepts. In such cases we need to think hard about the relationship between the thing we can measure and the thing we want to know.
Measurement reliability: describes how similar our measurement would be if we did it again. Because both the proportion of hybrid seed (out of eight) and pollinator visits in an average 15-minute observation session are associated with sampling error, they are not fully reliable. So we should remember that these are imperfect estimates of the thing we want to know. We could make them more reliable by e.g. watching longer or sequencing more plants, but this costs time and money. Other phenotypes, like petal area, can be measured with greater precision.
Internal validity
Internal validity describes how well a study design lets us draw the specific causal conclusion we’re actually after, given the individuals and conditions used in this study. A study has high internal validity when we can be confident that the effect we observe is due to the explanatory variable we care about, and not to some other factor that happens to be tangled up with it.
In making our RILs we did our best to mix up traits, but some trait correlations persisted. These correlations are a threat to internal validity – e.g. are associations between petal area and hybridization (Figure 2 A) due to petal area itself, or some correlated trait like anther-stigma distance (Figure 2 B), which is also associated with the proportion of hybrid seeds (Figure 2 C)? Luckily there are some paths we can walk down to assess how serious this threat to internal validity is.
External validity
External validity describes how far our results generalize beyond this study. We replicated this study across four field sites to see how consistent our results were across sites. We believe this study is broadly valid for the xantiana-parviflora system. However, it may not apply to other species such as those that are pollinated by wind or flies, or where nectar rewards are more important than floral display, etc.
This is important
Do not let perfection be the enemy of good
It is important to design a scientific study that can answer your question. You should think long and hard about this. But, also know that no study is the last word, all research decisions include trade-offs, and progress is better than nothing. So, think hard about making the best decision you can, and then ask if it is “good enough,” rather than aiming for the “perfect” experimental design.