为何使用dplyr::mutate_all后调用dplyr::summarise_all会报错?
Hey there! Let's break down why this is happening and fix it quickly.
The Root Cause
The issue here is that the scale() function returns a matrix (even for single columns) instead of a regular numeric vector. When you run mutate_all(scale), every column in your data frame gets converted to a matrix column. While each individual step works fine—mutate_all(scale) creates the matrix columns without issue, and summarise_all(mean) handles regular vectors perfectly—combining them trips up dplyr (especially in older versions) because it doesn't handle matrix columns smoothly during the summarization step.
Solutions
Fix 1: Convert scale results to vectors immediately
Wrap scale() with as.vector() to turn those matrix columns back into regular numeric vectors before summarizing. This gives dplyr the column type it expects:
mtcars %>% dplyr::mutate_all(~ as.vector(scale(.))) %>% dplyr::summarise_all(mean)
You'll notice the result has all zeros—this makes sense, since scaling centers each variable to have a mean of 0!
Fix 2: Use modern dplyr syntax (dplyr 1.0.0+)
If you're using a newer version of dplyr, across() is the recommended replacement for the _all() functions. You can handle scaling and vector conversion in a clean, future-proof way:
mtcars %>% dplyr::mutate(dplyr::across(everything(), ~ as.vector(scale(.)))) %>% dplyr::summarise(dplyr::across(everything(), mean))
Quick Verification
To confirm the matrix issue, run this to check the column types after mutation:
mtcars %>% dplyr::mutate_all(scale) %>% str()
You'll see each column is listed as num [1:32, 1]—that's a single-column matrix. Converting to a vector fixes this, letting summarise_all work as expected.
内容的提问来源于stack exchange,提问作者user3537951

