sparklyr无gather函数,如何实现类似宽表转长表操作?
gather with pivot_longer or Spark's stack Function Since gather() is part of tidyr and isn't directly available in sparklyr, you have a couple of straightforward alternatives to achieve the same long-format transformation you need:
1. Use dplyr::pivot_longer() (Recommended for Modern Sparklyr Versions)
Sparklyr now integrates seamlessly with dplyr's modern reshaping functions, and pivot_longer() is the official, updated replacement for gather(). It works almost identically to your original local tibble code:
# First, connect to Spark and copy your local tibble to the Spark cluster library(sparklyr) sc <- spark_connect(master = "local") a_spark <- copy_to(sc, a) # Perform the long-format transformation b_spark <- a_spark %>% pivot_longer( cols = -c(id, attribute1), # Keep id and attribute1 as-is; reshape all other columns names_to = "type_data", # Name for the new key column (matches your original `key` argument) values_to = "value_data" # Name for the new value column (matches your original `value` argument) ) # Optional: Bring the transformed data back to a local tibble b <- collect(b_spark)
This will produce exactly the structure you're expecting: columns value, average, upper_bound, and lower_bound will be stacked into type_data (the key) and value_data (the corresponding value), with id and attribute1 repeated for each row.
2. Use Spark SQL's stack() Function (For Older Sparklyr Versions)
If you're working with an older sparklyr version where pivot_longer() isn't supported, you can use Spark's built-in stack() function combined with dplyr:
b_spark <- a_spark %>% mutate( stacked = stack( list( value = value, average = average, upper_bound = upper_bound, lower_bound = lower_bound ) ) ) %>% select(id, attribute1, type_data = stacked.key, value_data = stacked.value) %>% unnest(cols = c(type_data, value_data))
The stack() function takes a named list of columns to reshape, creates a struct column with key and value fields, and we then unnest and rename these fields to match your desired output.
Quick Note on sdf_pivot
You mentioned hearing about sdf_pivot—it's important to clarify that this function is designed for widening data (similar to tidyr's spread() or pivot_wider()), not lengthening. It's not the right tool for replicating gather() behavior, so stick to the methods above instead.
内容的提问来源于stack exchange,提问作者RPisco

