dplyr::pull、purrr::pluck与magrittr::extract2的区别及适用场景
Great question! Let's break down these three functions—they all let you pull a single vector from a data frame (and beyond, in some cases) but have distinct strengths and ideal use cases, especially when working within the tidyverse ecosystem.
Core Differences, Advantages & Use Cases
1. dplyr::pull(): The Tidyverse Data Frame Workhorse
This is the go-to if you're already working in a dplyr pipeline. It's purpose-built for extracting columns from data frames (and tibbles) with tidyverse-friendly features that the other two don't offer:
- Tidy selection support: You can use helpers like
contains(),starts_with(), orends_with()to target columns without typing the exact name. - Vector naming: Use the
nameargument to set the names of your output vector using another column from the data frame. - Explicit error handling: If you reference a column that doesn't exist,
pull()throws a clear error instead of returningNULL(like the other two).
Best for: When you're wrapping up a dplyr workflow (filtering, mutating, grouping) and need to extract a column as a vector. It plays nicely with other tidyverse functions and makes your code more readable if you're already using dplyr syntax.
Example code:
# Extract by exact column name mtcars %>% mutate(wt_to_hp = wt/hp) %>% pull(wt_to_hp) # Extract by position (1 = first column, mpg) mtcars %>% pull(1) # Use tidy selector to grab a column matching a pattern mtcars %>% pull(contains("t")) # Returns the `wt` column # Name the output vector using values from another column mtcars %>% pull(wt, name = hp) # Vector of wt values, named with hp values
2. magrittr::extract2(): The Minimal [[ Pipe Wrapper
Think of this as a pipe-friendly version of base R's [[ operator. It does one thing and does it simply: extracts a single element (column, in data frame terms) using a name or position, just like df[[col]].
Advantages:
- Lightweight and consistent with other magrittr extract functions (like
extract()which is the pipe version of[). - No extra bells and whistles—perfect if you want a direct, base-R-equivalent operation in a pipe.
Best for: When you're using magrittr pipes but don't need the extra features of pull() or pluck(). It's a clean, minimal choice for straightforward column extraction.
Example code:
# Exact equivalent to mtcars[["wt"]] in pipe form mtcars %>% extract2("wt") # Extract by position mtcars %>% extract2(3) # Returns the `disp` column
3. purrr::pluck(): The Universal Nested Structure Extractor
This is the most flexible of the three—it's not limited to data frames. pluck() is designed to extract elements from any nested structure: lists, nested data frames, lists of lists, etc.
Key strengths:
- Multi-level extraction: You can pass multiple indices to dig into nested structures in one step (e.g., extract a column from a nested data frame inside a list).
- Works with all R objects: Not just data frames—use it to pull values from lists, S3 objects, or even nested tibbles.
- Graceful failure: Like
extract2()and[[, it returnsNULLif the element doesn't exist (instead of throwing an error).
Best for: When you're dealing with nested data (common in tasks like JSON parsing, grouped data with nest(), or list columns) and need to extract elements across multiple levels. It's overkill for simple data frame column pulls, but indispensable for complex nested structures.
Example code:
# Basic data frame column extraction (same as extract2() here) mtcars %>% mutate(wt_to_hp = wt/hp) %>% pluck("wt_to_hp") # Extract from a nested list nested_data <- list( metadata = list(year = 2024, source = "mtcars"), data = mtcars ) nested_data %>% pluck("metadata", "year") # Returns 2024 # Extract from a nested data frame mtcars %>% group_by(cyl) %>% nest() %>% pluck("data", 2, "hp") # Pulls hp column from the 6-cylinder group
Quick Cheat Sheet
| Function | Ideal Scenario | Unique Perks |
|---|---|---|
dplyr::pull() | Tidyverse data frame pipelines | Tidy selection, vector naming, clear errors |
magrittr::extract2() | Simple pipe-based [[ replacement | Magrittr style consistency, no frills |
purrr::pluck() | Nested lists/data frames | Multi-level extraction, works on all R objects |
内容的提问来源于stack exchange,提问作者crazybilly

