Jumping straight to the model might make us miss something interesting…
What Makes a Plot Interesting?
Prediction Edition
Sara Stoudt
Bucknell University
How do you go about finding something “interesting” in a new dataset? I recommend taking the time to explore. This exploration process requires actually looking at the data, making plots, and calculating basic summaries to start building some intuition about what information we have access to and what questions we might be able to ask and answer with it.
I’m in good company with this approach. Statistician John W. Tukey coined and popularized this process of exploratory data analysis (EDA) to contrast confirmatory data analysis (CDA). In CDA formal statistical analyses test for statistical significance, in other words deeming patterns we see in our explorations, meaningful in some way. However, Tukey argued that statisticians should avoid jumping into the analysis prematurely, and instead, spend some more time getting to know the data. This exploratory process can help us form stronger hypotheses that then can be formally tested later on. Being more familiar with a dataset also gives us more insight into what statistical tools may be appropriate for further investigating our hypotheses.
This exploration process becomes second nature to those who have a lot of experience with data, but the process can be intimidating for people just getting started with data. What does it mean to “explore”? How should we get started? In this post, I will try to formalize some of that intuition by unpacking what makes a graph “interesting”. Now, a graph can be interesting for a variety of reasons. In this post, I will focus on “interesting” in the context of building a prediction model. In a future post we’ll talk a little bit more about what makes a graph “interesting” from more of a narrative or storytelling point of view.
Let’s start exploring data appropriate for a simple scatterplot. We have a quantitative response variable $Y$ that we are interested in predicting with another quantitative explanatory variable $X$. Which of the available $X$s seem most interesting? Well, $X2$ seems more interesting than $X1$ because there seems to be a tighter relationship to $Y$. But what about $X3$? If we are too locked in to a particular modeling approach early on, perhaps linear regression here, we might disregard $X3$ even though it has a much tighter relationship to $Y$ than $X2$. Or if we just jump straight to fitting lines to points, we might miss that $X3$ has a curved relationship to $Y$ all together.
The exploration process gives us a little more time to get creative. Could we transform $X3$ to make the relationship with $Y$ more linear? Or could we learn about non-linear models that might work better than linear regression here? Those questions can lead us to some further exploration.

Now we’ll stick with a quantitative response $Y$, but suppose we only have categorical variables to choose from for our explanatory $X$. What should we be paying attention to in the (Tukey-invented) boxplots below to help us find something interesting? Let’s think about our goal. We want to predict $Y$, so knowing that $X$ is in category A vs. B should give us information about what values of $Y$ we expect to see. $X3$ doesn’t give us a lot to go on. Whether $X3$ takes the value A or B doesn’t help differentiate possible $Y$s. $X2$, on the other hand, looks a bit more interesting. It seems like $B$ values are associated with smaller values of $Y$. However, middle values of $Y$ could just as easily be connected to an A or B value of $X2$, so that variable only gets us so far.
What about $X1$? These boxplots don’t overlap at all, and that’s pretty interesting. If we know the value of $X1$ is B, then we expect $Y$ to be small, but if $X1$ is A, then we expect $Y$ to be large. This is what statisticians mean when they say they are looking for covariates that help “separate”; boxplots that are interesting for prediction don’t have much overlap across levels of the predictor variable. The same intuition about separation holds for predicting a categorical response $Y$ with a quantitative explanatory $X$.

But let’s complicate this intuition a bit. What about $X4$? Well, values of A seem interesting, but B and C, not so much. If we just threw $X4$ into a model we might disregard it as not being useful in predicting $Y$. But the picture gives us an idea of something further to explore. What if we adapt $X4$ to be a 1 if $X4$ is A and 0 otherwise? That will at least give us some predictive ability for smaller $Y$ values. Once again, slowing down and looking at the data gave us some further insight that we might have missed otherwise.

Now that we’ve built some intuition about what makes scatterplots and side-by-side boxplots interesting for predicting quantitative responses, let’s try the scenario that I find the most challenging to build intuition around: predicting a categorical response variable $Y$ with a categorical explanatory variable $X$.
A bar chart is fairly straightforward to interpret, whether it’s on a count or proportion scale. However, once we try to add a second categorical variable to the mix, we have to be a little more careful in our interpretation.
Below, each bar represents a level of the explanatory variable. Each bar is split into two (color/white) representing the two levels of the response variable. $Y$ is about equally spread across A and B of $X2$, and $X2$ is fairly balanced between A and B. However, one level of $Y$ is much more common in level A of $X1$ even though level A is less common in $X1$ itself. What does that tell us? Which $X$ is most interesting for predicting Y? Well frankly, it’s hard to tell. We aren’t really comparing apples to apples.

What if we normalize the heights of each level of $X1$ (the proportions within a level of $X$ add to 1)? Now we can see that the same proportion of each level of $X2$ corresponds to the colored level of $Y$. Similarly, we confirm that a higher proportion of A values correspond to the colored level of $Y$. Now it’s a little easier to tell that $X1$ gives us more information about $Y$ than $X2$. There is more “discrepancy” between levels of $Y$ within values of $X1$ compared to $X2$.

It’s still not a perfect comparison (we lose information about the relative split across the levels of the explanatory variable), but it will help us calibrate our sense of interesting-ness for prediction. Now, are you ready for a challenge?
Below is what is called a mosaic plot (it’s like a bar chart that accounts for more than two categorical variables) of passengers on the Titanic. We have information about their gender and class of ticket. Dark grey represents death while light grey represents survival. Which categorical predictor is most interesting for predicting survival status?

Where do you see the biggest differential in survival status? It might help to see what would happen if neither class nor gender were predictive of survival status. Below, everything is split proportionally as if survival was independent of gender and class. What part of the mosaic plot above is most different compared to the “no difference” plot below?

To me, it looks like gender is most predictive (women are much more likely to survive than men) although first class passengers also seem more likely to survive than those in other classes. But the great thing about statistics is that we don’t have to stop there. We could do some confirmatory data analysis to confirm (or overturn) my hypothesis. However, jumping straight to the model might make us miss something interesting. For example, it is interesting that the difference between second and third class survival for men doesn’t seem that different, but for women, there is a large decline in survival between second and third class. That suggests an intersection effect which we might not think to include otherwise.
It takes practice to make and interpret graphs and build intuition about what makes a plot interesting. Pick a dataset (there are a ton of interesting ones here) and explore your way towards a strong prediction model. Next time we’ll talk more about other ways graphs can be “interesting,” no models needed!
