如何用同国家非空值填充PySpark DataFrame中Continent列的空值?
PySpark 填充同一国家缺失的 Continent 字段
你的需求是用同一国家内已有的非空 Continent 值填充该国家内的空值,这里提供两种实用方法:
方法一:窗口函数(推荐)
利用窗口函数按国家分组,直接在原数据中提取非空值填充空记录,无需额外关联操作:
from pyspark.sql import Window from pyspark.sql.functions import first, col # 定义窗口:按Country分组,覆盖整个分组的行范围 window_spec = Window.partitionBy("Country").orderBy("Code").rowsBetween(Window.unboundedPreceding, Window.unboundedFollowing) # 用第一个非空的Continent值填充同国家内的空值 filled_df = df.withColumn( "Continent", first(col("Continent"), ignoreNulls=True).over(window_spec) ) # 查看结果 filled_df.show(truncate=False)
说明:
partitionBy("Country"):确保仅在同一国家的记录范围内处理first(..., ignoreNulls=True):跳过空值,获取分组内第一个有效的Continent值,也可以用last()函数,效果一致- 窗口范围设置为整个分组,保证能获取到该国家所有记录中的非空值
方法二:分组映射+关联
先构建国家与大洲的映射表,再通过左关联填充空值:
from pyspark.sql.functions import col # 生成国家-大洲映射表(过滤空值并去重) continent_map = df.filter(col("Continent").isNotNull()).select("Country", "Continent").distinct() # 左关联原表,用映射表的值填充空值 filled_df = df.join(continent_map, on="Country", how="left") \ .withColumn("Continent", col("Continent_y")) \ .drop("Continent_x", "Continent_y") filled_df.show(truncate=False)
说明:
- 映射表仅保留有有效大洲值的国家记录,去重后避免重复关联
- 左关联保证原表所有记录都保留,没有有效大洲值的国家(如France)会继续保留Null
两种方法都能得到你预期的输出结果,窗口函数方法更高效,适合处理大数据量场景。
内容的提问来源于stack exchange,提问作者Mohammad
相关产品推荐
相关产品推荐

