如何获取PySpark DataFrame中Vehicle列为空的Production列去重值列表
实现PySpark DataFrame按条件提取去重值并转为列表
需求说明
从给定的PySpark DataFrame中,筛选出Vehicle列为空的行,提取这些行中Production列的所有去重值,最终以字符串列表形式返回。
代码实现
首先创建测试用的DataFrame(模拟你的数据):
from pyspark.sql import SparkSession from pyspark.sql.functions import col # 初始化SparkSession spark = SparkSession.builder.appName("ExtractProduction").getOrCreate() # 创建测试DataFrame data = [ (1, "1234", 5678), (2, None, 1254), (3, None, 4567), (4, None, 4567) ] df = spark.createDataFrame(data, ["rowNum", "Vehicle", "Production"])
然后执行筛选、去重、转列表的操作:
# 筛选Vehicle为空的行,提取Production去重值,转为字符串列表 production_list = [str(row.Production) for row in df.filter(col("Vehicle").isNull()).select("Production").distinct().collect()] print(f"production list={production_list}")
代码解释
df.filter(col("Vehicle").isNull()):筛选出Vehicle列值为null的行.select("Production"):只保留Production列.distinct():对Production列的数值去重.collect():将分布式的DataFrame数据收集到本地,返回Row对象组成的列表- 列表推导式
[str(row.Production) ...]:将每个Row中的Production值转为字符串,组成最终的列表
运行后会输出:
production list=['1254', '4567']
内容的提问来源于stack exchange,提问作者karthik
相关产品推荐
相关产品推荐

