You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

dplyr中pull()与select()的区别:除单变量选择外还有其他差异吗?

dplyr::pull() vs dplyr::select(): Key Differences Beyond Variable Count

Great question! You’re spot-on that these two functions look similar at first glance—both help you pick out variables from a data frame—but there are critical differences beyond just pull() being limited to one variable. Let’s break them down:

1. Return Type: Data Frame vs. Vector

This is the biggest distinction:

  • select() always returns a tibble/data frame, even when you only pick one column. The result keeps the table structure, including column names.
  • pull() extracts the column as a raw vector (no table wrapper). It’s equivalent to using df$col or df[[col]], but fits cleaner in pipe workflows.

Example code to illustrate:

library(dplyr)

# select() returns a tibble
mtcars %>% select(mpg)
#> # A tibble: 32 × 1
#>     mpg
#>   <dbl>
#> 1  21  
#> 2  21  
#> 3  22.8
#> # … with 29 more rows

# pull() returns a numeric vector
mtcars %>% pull(mpg)
#>  [1] 21.0 21.0 22.8 21.4 18.7 18.1 14.3 24.4 22.8 19.2 17.8 16.4 17.3 15.2 10.4
#> [16] 10.4 14.7 32.4 30.4 33.9 21.5 15.5 15.2 13.3 19.2 27.3 26.0 30.4 15.8 19.7
#> [31] 15.0 21.4

2. Extra Flexibility in pull()

pull() has tricks select() doesn’t offer:

  • You can use negative indices to grab columns from the end (e.g., pull(-1) gets the last column of the data frame).
  • You can assign names to the returned vector using the name argument—this lets you turn one column into the names of another:
# Use cyl values as names for the mpg vector
mtcars %>% pull(mpg, name = cyl)
#>  6  6  4  6  8  6  8  4  4  6  6  8  8  8  8  8  8  4  4  4  4  8  8  8  8  4
#> 21.0 21.0 22.8 21.4 18.7 18.1 14.3 24.4 22.8 19.2 17.8 16.4 17.3 15.2 10.4 10.4 14.7 32.4 30.4 33.9 21.5 15.5 15.2 13.3 19.2 27.3
#>  4  4  8  6  8  4
#> 26.0 30.4 15.8 19.7 15.0 21.4

3. Use Case in Workflows

Choose based on what you need next:

  • Use select() when you want to keep working with a data frame (e.g., filtering, mutating, or joining after selecting columns). It preserves the tabular structure for seamless dplyr chaining.
  • Use pull() when you need raw vector data (e.g., calculating a mean with mean(), feeding into a plot’s x/y axis, or using base R functions that expect vectors instead of data frames).

Quick Summary

Think of select() as "keeping columns in the table" and pull() as "taking a column out of the table to use as a standalone vector."

内容的提问来源于stack exchange,提问作者Evan O.

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 04:00:56