What Makes a Plot Interesting?

Jumping straight to the model might make us miss something interesting…

What Makes a Plot Interesting?

Prediction Edition

Sara Stoudt
Bucknell University

How do you go about finding something “interesting” in a new dataset? I recommend taking the time to explore. This exploration process requires actually looking at the data, making plots, and calculating basic summaries to start building some intuition about what information we have access to and what questions we might be able to ask and answer with it.

I’m in good company with this approach. Statistician John W. Tukey coined and popularized this process of exploratory data analysis (EDA) to contrast confirmatory data analysis (CDA). In CDA formal statistical analyses test for statistical significance, in other words deeming patterns we see in our explorations, meaningful in some way. However, Tukey argued that statisticians should avoid jumping into the analysis prematurely, and instead, spend some more time getting to know the data. This exploratory process can help us form stronger hypotheses that then can be formally tested later on. Being more familiar with a dataset also gives us more insight into what statistical tools may be appropriate for further investigating our hypotheses.

This exploration process becomes second nature to those who have a lot of experience with data, but the process can be intimidating for people just getting started with data. What does it mean to “explore”? How should we get started? In this post, I will try to formalize some of that intuition by unpacking what makes a graph “interesting”. Now, a graph can be interesting for a variety of reasons. In this post, I will focus on “interesting” in the context of building a prediction model. In a future post we’ll talk a little bit more about what makes a graph “interesting” from more of a narrative or storytelling point of view.

Let’s start exploring data appropriate for a simple scatterplot. We have a quantitative response variable $Y$ that we are interested in predicting with another quantitative explanatory variable $X$. Which of the available $X$s seem most interesting? Well, $X2$ seems more interesting than $X1$ because there seems to be a tighter relationship to $Y$. But what about $X3$? If we are too locked in to a particular modeling approach early on, perhaps linear regression here, we might disregard $X3$ even though it has a much tighter relationship to $Y$ than $X2$. Or if we just jump straight to fitting lines to points, we might miss that $X3$ has a curved relationship to $Y$ all together.

The exploration process gives us a little more time to get creative. Could we transform $X3$ to make the relationship with $Y$ more linear? Or could we learn about non-linear models that might work better than linear regression here? Those questions can lead us to some further exploration.

Sketches of three scatterplots with a common y-axis. Each sketch shows a relationship with a potential quantitative X. The first is a positive, linear relationship that is rather weak. The second is a stronger positive, linear relationship. The third is an even stronger positive relationship, but it is non-linear, plateauing after a certain point.
When we explore relationships between two quantitative variables we are looking for narrow bands. $X2$ has a narrower relationship to $Y$ than $X1$, but $X3$ has a narrower relationship to $Y$ than $X2$ (but that relationship is non-linear).

Now we’ll stick with a quantitative response $Y$, but suppose we only have categorical variables to choose from for our explanatory $X$. What should we be paying attention to in the (Tukey-invented) boxplots below to help us find something interesting? Let’s think about our goal. We want to predict $Y$, so knowing that $X$ is in category A vs. B should give us information about what values of $Y$ we expect to see. $X3$ doesn’t give us a lot to go on. Whether $X3$ takes the value A or B doesn’t help differentiate possible $Y$s. $X2$, on the other hand, looks a bit more interesting. It seems like $B$ values are associated with smaller values of $Y$. However, middle values of $Y$ could just as easily be connected to an A or B value of $X2$, so that variable only gets us so far.

What about $X1$? These boxplots don’t overlap at all, and that’s pretty interesting. If we know the value of $X1$ is B, then we expect $Y$ to be small, but if $X1$ is A, then we expect $Y$ to be large. This is what statisticians mean when they say they are looking for covariates that help “separate”; boxplots that are interesting for prediction don’t have much overlap across levels of the predictor variable. The same intuition about separation holds for predicting a categorical response $Y$ with a quantitative explanatory $X$.

Sketches of three side-by-side boxplots (two per plot) with a common y-axis. Each sketch shows a relationship with a potential categorical X. The first has boxplots completely separated. The second has boxplots that partially overlap in range. The third has boxplots that are almost identical.
When we explore relationships between a quantitative response variable and a categorical explanatory variable we are looking for a lack of overlap in side-by-side boxplots. $X2$ has less overlap in $Y$ across its two levels than $X3$, but $X1$ does not have any overlap at all.

But let’s complicate this intuition a bit. What about $X4$? Well, values of A seem interesting, but B and C, not so much. If we just threw $X4$ into a model we might disregard it as not being useful in predicting $Y$. But the picture gives us an idea of something further to explore. What if we adapt $X4$ to be a 1 if $X4$ is A and 0 otherwise? That will at least give us some predictive ability for smaller $Y$ values. Once again, slowing down and looking at the data gave us some further insight that we might have missed otherwise.

Sketches of a side-by-side boxplot with three levels. Two levels completely overlap in terms of the Y value range while one does not overlap the other two at all.
Levels B and C are associated with similar ranges of $Y$ while level A is associated with lower values of $Y$. This plot motivates us to try to collapse levels B and C into one to better predict $Y$ from $X4$.

Now that we’ve built some intuition about what makes scatterplots and side-by-side boxplots interesting for predicting quantitative responses, let’s try the scenario that I find the most challenging to build intuition around: predicting a categorical response variable $Y$ with a categorical explanatory variable $X$.

A bar chart is fairly straightforward to interpret, whether it’s on a count or proportion scale. However, once we try to add a second categorical variable to the mix, we have to be a little more careful in our interpretation.

Below, each bar represents a level of the explanatory variable. Each bar is split into two (color/white) representing the two levels of the response variable. $Y$ is about equally spread across A and B of $X2$, and $X2$ is fairly balanced between A and B. However, one level of $Y$ is much more common in level A of $X1$ even though level A is less common in $X1$ itself. What does that tell us? Which $X$ is most interesting for predicting Y? Well frankly, it’s hard to tell. We aren’t really comparing apples to apples.

Sketches of two clustered bar charts of two levels each. The bars represent levels of potential categorical explanatory variables while each bar is split by color representing the two levels of the categorical response Y. The first chart shows one level of Y being much more common in one level of X than the other while in the second chart each level of Y is fairly evenly spread across the two levels of X.
When we explore relationships between two categorical variables we are looking for differences in the distribution of the response variable across the levels of the explanatory variable. Each level of $X2$ has roughly the same distribution of Y. The A level of $X1$ has many more instances of one $Y$ level than the B level. However, the differences in how many instances of each explanatory variable level there are make it hard to be sure that difference is meaningful.

What if we normalize the heights of each level of $X1$ (the proportions within a level of $X$ add to 1)? Now we can see that the same proportion of each level of $X2$ corresponds to the colored level of $Y$. Similarly, we confirm that a higher proportion of A values correspond to the colored level of $Y$. Now it’s a little easier to tell that $X1$ gives us more information about $Y$ than $X2$. There is more “discrepancy” between levels of $Y$ within values of $X1$ compared to $X2$.

Sketches of two clustered bar charts of two levels each now normalized so the values within a level of the explanatory variable sum to one. The bars represent levels of potential categorical explanatory variables, and they still split by color representing the two levels of the categorical response Y. The first chart shows one level of Y appearing at a higher proportion in one level of X than the other while in the second chart each level of Y is fairly evenly spread across the two levels of X.
Normalizing the clustered bar charts in this way helps us definitely say that a higher proportion of $X1$ data whose values is A is the colored level of $Y$ and confirm that $X2$ doesn’t seem as interesting for predicting values of $Y$.

It’s still not a perfect comparison (we lose information about the relative split across the levels of the explanatory variable), but it will help us calibrate our sense of interesting-ness for prediction. Now, are you ready for a challenge?

Below is what is called a mosaic plot (it’s like a bar chart that accounts for more than two categorical variables) of passengers on the Titanic. We have information about their gender and class of ticket. Dark grey represents death while light grey represents survival. Which categorical predictor is most interesting for predicting survival status?

A mosaic plot of Titanic passengers with two categorical predictors (gender and class) and whether or not they survived. The biggest discrepancies are between female and male passengers (with females more likely to survive) and across class (with first class more likely to survive than other classes)
This mosaic plot of Titanic passengers encodes two categorical predictors (gender and class) and whether or not they survived (dark grey = death, light grey = survival). We are looking for interesting discrepancies in shades of grey across the two predictors.

Where do you see the biggest differential in survival status? It might help to see what would happen if neither class nor gender were predictive of survival status. Below, everything is split proportionally as if survival was independent of gender and class. What part of the mosaic plot above is most different compared to the “no difference” plot below?

A fictional mosaic plot of Titanic passengers with two categorical predictors (gender and class) and whether or not they survived. In this counterfactual scenario, there are no discrepancies based on gender or class.
Seeing what a mosaic plot would look like if there was no difference in survival (dark grey = death, light grey = survival) This mosaic plot of titanic passengers encodes two categorical predictors (gender and class) and whether or not they survived (dark grey = death, light grey = survival). We are looking for interesting discrepancies in shades of grey across the two predictors.

To me, it looks like gender is most predictive (women are much more likely to survive than men) although first class passengers also seem more likely to survive than those in other classes. But the great thing about statistics is that we don’t have to stop there. We could do some confirmatory data analysis to confirm (or overturn) my hypothesis. However, jumping straight to the model might make us miss something interesting. For example, it is interesting that the difference between second and third class survival for men doesn’t seem that different, but for women, there is a large decline in survival between second and third class. That suggests an intersection effect which we might not think to include otherwise.

It takes practice to make and interpret graphs and build intuition about what makes a plot interesting. Pick a dataset (there are a ton of interesting ones here) and explore your way towards a strong prediction model. Next time we’ll talk more about other ways graphs can be “interesting,” no models needed!

Leave a Reply

Your email address will not be published. Required fields are marked *

HTML tags are not allowed.

86,900 Spambots Blocked by Simple Comments