Data management with tidyverse

The tidyverse is an opinionated collection of R packages designed for data science. All packages share an underlying design philosophy, grammar, and data structures.

To learn more about the tidyverse, visit https://tidyverse.org/learn/

Work through the R for Data Science book: https://r4ds.hadley.nz/

Overview

tidyverse is a collection of R packages:

  • dplyr – for data manipulation
  • tidyr – for making data “tidy”
  • readr – for efficiently reading data
  • stringr – for working with strings
  • ggplot2 – for data visualization
  • forcats – utility functions for working with factor data
  • lubridate – for working with date/time data
  • purrr – tools for working with functions and vectors
  • tibble – a ‘modern reimagining’ of the data frame

Round-robin tour of some highlights

We will make use of the penguins data frame for this set of exercises

dplyr

The main goal is to make it easier and faster to deal with data and data manipulations

  • Over a hundred functions – lots to learn!

  • A few key categories of functions:

  • mutate amd mutate_*: add columns to a dataframe

  • filter and slice: grab rows

  • group_by: group dataframe for grouped operations

  • summarize: for creating summaries (within groups if appropriate)

  • *_join: for merging dataframes

Grouped calculations with summarize

Or we can group by multiple categories

# Group by species and island
penguins |> 
  group_by(species, island) |> 
  summarize(mean_billLength = mean(bill_length_mm, na.rm = T))

Combining tibbles with join functions

tidyr

Main goal is to facilitate tidy data operations:

  • switching from a wide to long format
  • separating and/or uniting information
  • nesting tibbles

Switching data from wide to long format

Uniting or separating information

Nesting grouped datasets

Crossing information to create tibbles

readr

Main goal is to read in data files

  • There are already base functions in R to do this, e.g. read.csv
  • readr provides alternatives that do some things better, e.g. read_csv
  • But there’s not a strictly “better” option – readr can be nice if are relying on other tidyverse functions

ggplot2

Main goal is to provide a consistent “grammar of graphics”

  • Now a highly used framework for plotting, so lots of extensions available

  • We will do a deeper dive into ggplot2 later on.

Combining components of the tidyverse for effective plotting

forcats

Designed for helping deal with categorical (factor) data

What if we wanted to arrange the species another way?

What if we also wanted the order within each species to be male – female

purrr

For facilitating functional programming in R.

  • The goal of functional programming is to heavily rely on functions to create useful software
  • Extremely powerful (), but can be hard to start thinking this way if you are not used to it

stringr

For working with string data

Applications in Ecology and Evolution

Combining these tools is an extremely powerful approach for writing compact, legible code

e.g. the following chunk is running 10,000 models simultaneously, and then plotting the results from all 10,000:

out_det <-
  crossing(
    # 1. p00 2. pG0 3. pL0 4. pGG 5. pSS 6. p0G 7. p0S 8. pLS 9. pGS 10. pLG
    nvec = "n0 = [0.1, 0.1, 0.1, 0.6, 0.1, 0.0, 0.0, 0.0, 0.0, 0.0];",
    # nvec = "n0 = [0.1, 0.1, 0, 0, 0.8, 0.0, 0.0, 0.0, 0.0, 0.0];",
    # nvec = "n0 = [0.2, 0, 0, 0.4, 0.4, 0.0, 0.0, 0.0, 0.0, 0.0];",
    rs = list(c(rG = 1, rS = 1.55)),
    ms = list(c(mG = 0.1, mS = 0.1)),
    ds = list(c(dG = 0.25, dS = 0.25)),
    cs = list(c(cG = 0.2, cS = 0.2)),
    mGG = crossing(mSS = seq(0.5,1.5, 0.01), mGG = seq(0.5,1.5, 0.01)) |> pull(mGG),
    mSS = crossing(mSS = seq(0.5,1.5, 0.01), mGG = seq(0.5,1.5, 0.01)) |> pull(mSS),
    time = 2000,
    fint = 50000, 
    intensity = 0.2) |> 
  rowwise() |> 
  mutate(mijs = list(c(mGG = mGG, mSG = 1, mGS = 1, mSS = mSS)),
         pvec = list(c(rs,ms,ds,cs, mijs)),
         pvec = paste0("p = (", vec_to_tup(pvec), ");")) |>
  mutate(julia_out = list(run_model_ffp(n = nvec, p = pvec, time = time, 
                                        fireint = fint, intensity = intensity,  ffp = 0,
                                        which_model = "fire_shrubstructured"))) 

plot_det <- 
  out_det |> 
  mutate(julia_out = list(julia_out[julia_out[,"timestamp"] > 1950,] )) |>
  mutate(pGrass = mean(rowSums(julia_out[,c("pG0(t)", "pGG(t)", "pGS(t)")])),
         pShrub = mean(rowSums(julia_out[,c("pL0(t)", "pLS(t)", "pLG(t)", "pSS(t)")])),
         pEmpty = mean(rowSums(julia_out[,c("p00(t)", "p0G(t)", "p0S(t)")]))) |>
  group_by(fint, intensity) |>
  group_by(mGG, mSS) |>
  summarize(mean_pGrass = mean(pGrass)) |>
  ggplot() + 
  geom_tile(aes(x = mGG, y = mSS, fill = mean_pGrass), color = NA) + 
  geom_hline(yintercept = 1, color = "white", linewidth = 1.5) + 
  geom_vline(xintercept = 1, color = "white", linewidth = 1.5) + 
  scale_fill_viridis_c(name = "Equilibrium grass frequency")+
  scale_x_continuous(expand = c(0,0))+
  scale_y_continuous(expand = c(0,0))+
  labs(x = latex2exp::TeX("$m_{gg}$"), 
       y=latex2exp::TeX("$m_{ss}$"), tag = "A. ")

##

Let’s do some exercises

Exercises

Lots of resources to learn more: