You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

sparklyr无gather函数,如何实现类似宽表转长表操作?

Solution for Sparklyr: Replace gather with pivot_longer or Spark's stack Function

Since gather() is part of tidyr and isn't directly available in sparklyr, you have a couple of straightforward alternatives to achieve the same long-format transformation you need:

Sparklyr now integrates seamlessly with dplyr's modern reshaping functions, and pivot_longer() is the official, updated replacement for gather(). It works almost identically to your original local tibble code:

# First, connect to Spark and copy your local tibble to the Spark cluster
library(sparklyr)
sc <- spark_connect(master = "local")
a_spark <- copy_to(sc, a)

# Perform the long-format transformation
b_spark <- a_spark %>%
  pivot_longer(
    cols = -c(id, attribute1),  # Keep id and attribute1 as-is; reshape all other columns
    names_to = "type_data",      # Name for the new key column (matches your original `key` argument)
    values_to = "value_data"     # Name for the new value column (matches your original `value` argument)
  )

# Optional: Bring the transformed data back to a local tibble
b <- collect(b_spark)

This will produce exactly the structure you're expecting: columns value, average, upper_bound, and lower_bound will be stacked into type_data (the key) and value_data (the corresponding value), with id and attribute1 repeated for each row.

2. Use Spark SQL's stack() Function (For Older Sparklyr Versions)

If you're working with an older sparklyr version where pivot_longer() isn't supported, you can use Spark's built-in stack() function combined with dplyr:

b_spark <- a_spark %>%
  mutate(
    stacked = stack(
      list(
        value = value,
        average = average,
        upper_bound = upper_bound,
        lower_bound = lower_bound
      )
    )
  ) %>%
  select(id, attribute1, type_data = stacked.key, value_data = stacked.value) %>%
  unnest(cols = c(type_data, value_data))

The stack() function takes a named list of columns to reshape, creates a struct column with key and value fields, and we then unnest and rename these fields to match your desired output.

Quick Note on sdf_pivot

You mentioned hearing about sdf_pivot—it's important to clarify that this function is designed for widening data (similar to tidyr's spread() or pivot_wider()), not lengthening. It's not the right tool for replicating gather() behavior, so stick to the methods above instead.

内容的提问来源于stack exchange,提问作者RPisco

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 09:21:53