You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用flatMap+distinct能否从RDD嵌套数组生成去重元素列表?

解决RDD数组元素拆分去重的方案

嘿,你已经找对核心方法啦!flatMap + distinct正是处理这个需求的标准操作,我来帮你把代码补全并解释清楚细节:

完整代码示例

# 假设你已经初始化好SparkContext(命名为sc)
rdd = sc.parallelize([
    ['Adventure', 'Animation', 'Children', 'Comedy', 'Fantasy'],
    ['Adventure', 'Children', 'Fantasy'],
    ['Comedy', 'Romance'],
    ['Comedy', 'Drama', 'Romance'],
    ['Comedy']
])

# 步骤1:将每个数组的元素拆分为单独行;步骤2:去重
unique_genres_rdd = rdd.flatMap(lambda x: x).distinct()

# 若需要转换为本地Python列表(小数据量场景适用)
unique_genres_list = unique_genres_rdd.collect()
print(unique_genres_list)

代码解释

  • flatMap(lambda x: x):这个操作会把RDD中的每个数组“拍平”——简单来说就是把数组里的每个元素单独提取出来,变成RDD的独立元素。比如原来的5个数组会被拆成14个单独的genre条目。
  • .distinct():对拆分后的所有元素进行去重,只保留每个genre的唯一实例。
  • collect():把分布式RDD中的数据拉取到本地Driver节点,转换成Python列表。注意:如果你的RDD数据量极大,不建议直接用collect(),避免内存溢出。

可选扩展:排序输出

如果希望最终的去重列表是按字母顺序排列的,可以在distinct()后添加sortBy操作:

sorted_unique_genres = rdd.flatMap(lambda x: x).distinct().sortBy(lambda x: x).collect()
print(sorted_unique_genres)

内容的提问来源于stack exchange,提问作者michael green

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 06:51:53