dplyr中pull()与select()的区别:除单变量选择外还有其他差异吗?
dplyr::pull() vs dplyr::select(): Key Differences Beyond Variable Count
Great question! You’re spot-on that these two functions look similar at first glance—both help you pick out variables from a data frame—but there are critical differences beyond just pull() being limited to one variable. Let’s break them down:
1. Return Type: Data Frame vs. Vector
This is the biggest distinction:
select()always returns a tibble/data frame, even when you only pick one column. The result keeps the table structure, including column names.pull()extracts the column as a raw vector (no table wrapper). It’s equivalent to usingdf$colordf[[col]], but fits cleaner in pipe workflows.
Example code to illustrate:
library(dplyr) # select() returns a tibble mtcars %>% select(mpg) #> # A tibble: 32 × 1 #> mpg #> <dbl> #> 1 21 #> 2 21 #> 3 22.8 #> # … with 29 more rows # pull() returns a numeric vector mtcars %>% pull(mpg) #> [1] 21.0 21.0 22.8 21.4 18.7 18.1 14.3 24.4 22.8 19.2 17.8 16.4 17.3 15.2 10.4 #> [16] 10.4 14.7 32.4 30.4 33.9 21.5 15.5 15.2 13.3 19.2 27.3 26.0 30.4 15.8 19.7 #> [31] 15.0 21.4
2. Extra Flexibility in pull()
pull() has tricks select() doesn’t offer:
- You can use negative indices to grab columns from the end (e.g.,
pull(-1)gets the last column of the data frame). - You can assign names to the returned vector using the
nameargument—this lets you turn one column into the names of another:
# Use cyl values as names for the mpg vector mtcars %>% pull(mpg, name = cyl) #> 6 6 4 6 8 6 8 4 4 6 6 8 8 8 8 8 8 4 4 4 4 8 8 8 8 4 #> 21.0 21.0 22.8 21.4 18.7 18.1 14.3 24.4 22.8 19.2 17.8 16.4 17.3 15.2 10.4 10.4 14.7 32.4 30.4 33.9 21.5 15.5 15.2 13.3 19.2 27.3 #> 4 4 8 6 8 4 #> 26.0 30.4 15.8 19.7 15.0 21.4
3. Use Case in Workflows
Choose based on what you need next:
- Use
select()when you want to keep working with a data frame (e.g., filtering, mutating, or joining after selecting columns). It preserves the tabular structure for seamless dplyr chaining. - Use
pull()when you need raw vector data (e.g., calculating a mean withmean(), feeding into a plot’s x/y axis, or using base R functions that expect vectors instead of data frames).
Quick Summary
Think of select() as "keeping columns in the table" and pull() as "taking a column out of the table to use as a standalone vector."
内容的提问来源于stack exchange,提问作者Evan O.
相关产品推荐
相关产品推荐

