Data visualization: theory and practice

We tend to make inferences about relationships between the objects that we see in ways that bear on our interpretation of graphical data, for example. Arrangements of points and lines on a page can encourage us—sometimes quite unconsciously—to make inferences about similarities, clustering, distinctions, and causal relationships that might or might not be there in the numbers. Sometimes these perceptual tendencies can be honestly harnessed to make our graphics more effective. At other times, they will tend to lead us astray, and must take care not to lean on them too much.

Anscombe’s quartet

Code
library(tidyverse)
library(glue)
library(gt)
# reshape anscombe data in tidy format
ans <- 
  anscombe |> 
  as_tibble() |> 
  pivot_longer(x1:y4) |> 
  mutate(dataset = case_when(str_detect(name, "1") ~ "dataset1",
                             str_detect(name, "2") ~ "dataset2",
                             str_detect(name, "3") ~ "dataset3",
                             str_detect(name, "4") ~ "dataset4")) |>
  mutate(xory = ifelse(str_detect(name, "x"), "x", "y")) |> 
  select(-name) |> 
  pivot_wider(names_from = xory, values_from = value, values_fn = list) |> 
  unnest(cols = c("x","y")) 

# Define a function to make labels for plots
labler <- function(rsq,pval) glue("R = {rsq}, p = {pval}")

# generate linear models, extract p-vals and rsq; generate labels
ans_sum <- 
  ans |> 
  group_by(dataset) |> 
  nest() |> 
  mutate(linmod = map(data, \(data)
                      lm(y~x, data = data)),
         linmod_s = map(linmod, summary),
         rsq = map_dbl(linmod_s, \(sum) round(sqrt(sum$r.squared),3)),
         pval = map_chr(linmod_s, \(sum) scales::pvalue(sum$coefficients[2,4])),
         label = map2_chr(rsq, pval, labler))
ans_plot <-
  ans |> 
  ggplot(aes(x = x, y = y))  +
  geom_point() +
  facet_wrap(.~dataset, nrow = 2) + 
  theme_bw()

ans_plot

Anscombe’s quartet

Code
ans_plot + 
  geom_smooth(method = "lm", se = F)

Anscombe’s quartet

Code

ans_plot + 
  geom_smooth(method = "lm", se = F) + 
  geom_text(inherit.aes = F, data = ans_sum,
            aes(x = Inf,y = -Inf, label = label),
            hjust = 1.1, vjust = -1) 

Illustrations like this demonstrate why it is worth looking at data. But that does not mean that looking at data is all one needs to do. Real datasets are messy, and while displaying them graphically is very useful, doing so presents problems of its own. As we will see below, there is considerable debate about what sort of visual work is most effective, when it can be superfluous, and how it can at times be misleading to researchers and audiences alike.

Plotting in R with ggplot2

ggplot2 is an R package for producing statistical, or data, graphics. Unlike most other graphics packages, ggplot2 has an underlying grammar, based on the Grammar of Graphics (Wilkinson 2005), that allows you to compose graphs by combining independent components.

— Hadley Wickham, Danielle Navarro, and Thomas Lin Pedersen, ggplot2: Elegant Graphics for Data Analysis (3e)

ggplot2 origin & logic

Wilkinson 2006 created the grammar of graphics to describe the fundamental features that underlie all statistical graphics.

The grammar of graphics is an answer to the question of what is a statistical graphic?

ggplot2 builds on Wilkinson’s grammar by focusing on the primacy of layers and adapting it for use in R.

In brief, the grammar tells us that a graphic maps the data to the aesthetic attributes (colour, shape, size) of geometric objects (points, lines, bars).

The plot may also include statistical transformations of the data and information about the plot’s coordinate system.

Facetting can be used to plot for different subsets of the data.

The combination of these independent components are what make up a graphic.

Anatomy of a graph

Quoted from ggplot2 book:

All plots are composed of the data, the information you want to visualise, and a mapping, the description of how the data’s variables are mapped to aesthetic attributes.

There are five mapping components:

Mapping component 1

  • A layer is a collection of geometric elements (*geoms) and statistical transformations (stats).

geoms represent what you actually see in the plot: points, lines, polygons, etc. stats summarise the data: for example, binning and counting observations to create a histogram, or fitting a linear model

Mapping component 2

  • Scales map values in the data space to values in the aesthetic space. This includes the use of colour, shape or size. Scales also draw the legend and axes, which make it possible to read the original data values from the plot (an inverse mapping).

Mapping component 3

  • A coord, or coordinate system, describes how data coordinates are mapped to the plane of the graphic. We normally use the Cartesian coordinate system, but a number of others are available, including polar coordinates and map projections.

Mapping component 4

  • A facet specifies how to break up and display subsets of data as small multiples.

Mapping component 5

  • A theme controls the finer points of display, like the font size and background colour.

This is a lot to take in - but its good to have in your back pocket.

For now, remember that Every ggplot2 plot has three key components:

  1. Data,

  2. A set of aesthetic mappings between variables in the data and visual properties, and

  3. At least one layer which describes how to render each observation. Layers are usually created with a geom function.

Some practical examples

We will work with the gapminder dataset in R

For simplicity, we’ll work with data for the 6 largest countries (by population) as of 2007.

Let’s make some plots

  1. Map Year to X-axis, Life Expectancy to Y-axis; add a “point” geometry

Let’s make some plots

  1. Map Year to X-axis, Life Expectancy to Y-axis with a point geometry, but also map country to color.

  1. Let’s try something different. Map per-capita GDP to X-axis, and Life Expectancy to Y-axis; show using a point geometry

  1. Year was nowhere represented. How can we include that?

Same aesthetics, different geometries

  1. Map country to color; Year to X-axis, Life Expectancy to Y-axis. Show points for each entry.

  1. Map country to color; Year to X-axis, Life Expectancy to Y-axis. Show lines for each entry.

  1. Map country to color; Year to X-axis, Life Expectancy to Y-axis. Show points and lines for each entry.

Note that sometimes, graphs have the right syntax but are meaningless

  • Map country to color, Year to x-axis, Life expectancy to Y-axis, and show using a boxplot geometry

See also Noam Chomsky’s “Colorless green ideas sleep furiously”

Same aesthetics and geometries, different scales

  1. Map country to color; Year to X-axis, Life Expectancy to Y-axis. Show points and lines for each entry. Use default scales.

  1. Map country to color; Year to X-axis, Life Expectancy to Y-axis. Show points and lines for each entry. Scale the colors manually.

  1. Map country to color; Year to X-axis, Life Expectancy to Y-axis. Show points and lines for each entry. Scale the colors using a pre-defined palette.

Use “small multiples” for bigger datasets

Use the full gapminder dataset to make a similar plot of life expectancy over time for all countries . . .

Use “small multiples” for bigger datasets

Simplify this by showing one small panel per continent

Your turn to practice

  • If you are new to R, consider working through the Data Carpentry’s ggplot2 lesson for ecologists

  • If you feel comfortable with the basics, go to the next slide for a ‘plotting challenge’ that combines various tidyverse packages, including ggplot2

Use the following code to import in a dataset of the birds recorded from the LSU Lakes Ebird Hotspot in 2023 and 2024:

ebird_out <- readRDS(url("https://gitlab.com/gklab/teaching/rr-f26/-/raw/main/weekly-activities/week-05/ebird-UniversityLakes.rds"))

Use this information to make the following plot:

In-class code

Code that we worked on in class…

ebird_out <- readRDS(url("https://gitlab.com/gklab/teaching/rr-f26/-/raw/main/weekly-activities/week-05/ebird-UniversityLakes.rds"))

ebird_out |> 
  unnest(ebird_out) |> 
  select(year, months, comName, sciName, howMany) |> 
  filter(year == 2024) |> 
  group_by(months, comName) |> 
  summarise(total_seen = sum(howMany),
            n_obsevers = n()) |> 
  mutate(avg_obs = total_seen/n_obsevers) |> 
  ggplot(aes(x = avg_obs, y = comName)) +
  geom_point() +
  facet_wrap(~ months, scales = "free")

# Filter the data set to show only the 20 most abundant species in each month
# Arrange the species by number of observations instead of alphabeticaly
# Modify the facet label to include total # of species seen that month.