Pipes in R

Most analysis workflows involve multiple steps, e.g.:

  • “Take this dataframe, create a new column by combining information from two columns, then take the log of the new column, and then tell me the mean value”

  • We’ll work on this with the penguins dataframe on the following slide

Approach 1

One approach looks something like this:

library(palmerpenguins)

penguins$bill_ratio <- penguins$bill_length_mm/penguins$bill_depth_mm

log_bill_ratio <- log(penguins$bill_ratio)

mean_log_bill_ratio <- mean(log_bill_ratio, na.rm = TRUE) # we need na.rm = TRUE to remove some of the NA values

# Finally this gets us the answer:
mean_log_bill_ratio
[1] 0.9393296

Approach 2

The second approach is to nest all the steps within one another:

mean(log(penguins$bill_length_mm/penguins$bill_depth_mm), na.rm = T)
[1] 0.9393296

Approach 3 (with pipes!)

Approach 1 worked, and was easy to follow… but if you had a lot of steps, it could get tedious

Approach 2 also worked for this example… but things get more complicated in real life, and such “nested” code can be quite hard to read!

  • Reading that code would get hard!
  • Pipes to the rescue:
(penguins$bill_length_mm/penguins$bill_depth_mm) |> 
  log() |> 
  mean(na.rm = T)
[1] 0.9393296